Category: OpenCV Live

  • 06/25/2026 – OpenCV Live

    Hey there. Hi there. Hello there, everybody. It is Thursday. It’s 9:00 a.m. You know what that means? It’s time once again for Open CV Live. We’re so happy to have you all here today. It’s going to be an educational episode with our friend Dennis Baldwin here from Drone Blocks. We’ll give you all a few minutes to get situated, get on all your various platforms. We’re live on Twitch, LinkedIn live video, YouTube, and now on live streaming on our Patreon, which is pretty exciting um to be able to bring this stuff straight to patrons. So, if you want to see OpenCV Live live without any of those pesky YouTube ads, you can become an OpenCV patron at patreon.com. I’ll talk a little bit more about that later in the episode. Great. got some nice stuff happening already over on uh YouTube and uh let’s see how see how our folks on Zoom are making it happen today. Indeed, we’ve got Carlos Carlos Nav says nice. Yes, Vamos Mexico uh 3-0 yesterday. A very exciting game in the World Cup. I hope that all of your teams are moving on to the from the group round, but if not, uh there’s always next time. They say it’s the hope that kills you, you know, but sometimes the hope is all we got. So, yeah, thanks. We’ve got uh Stefan SA checking in from Vancouver, Canada. So, as you join us today, wherever you’re watching, wherever you’re joining us from, uh let us know where you’re joining us from. We’d love to see the global OpenCV community chiming in on these Thursday live streams. I’m coming to you from beautiful, far too sunny for my pasty white skin Tana, Mexico. However, uh it was it has been super great to be able to be here in Mexico and watch the World Cup with uh the the people of Mexico and the excitement is so the excite it’s always you know with with these countries that are very very into uh you know soccer. It’s the it’s always excitement tempered with realism. Like they know that they’re not favorites, right? But and they know that they’re very hard on the team, but it’s just because they care so much and it’s it’s been it’s been really fun. You you you you clearly have not seen the ad where it says uh American you cannot smell spell American without I can. [laughter] Um well I I think they’re going to win. You have not seen the miracle ad, have you? No, you have not been seeing all that. I haven’t. But in Mexico it’s uh it’s Cis. Yes, we can. Um, we’ve got uh Madagascar, we’ve got Houston, uh, Edward Bragg says, “Uh, Merry Mryland.” I assume that’s Maryland, but I like that. Uh, we’ve got, uh, Axel Tracy says, “Kickoff time for the Sakaroos in 10 hours.” Hello from Australia. Good luck to the Aussies. We’ve got Petro from Saudi Arabia. We’ve got quite a few of the AI assistants today. We’ve got Doug chiming in from uh Pmpano Beach. That’s a that’s a fun name to say. And we’ve got uh Carlos from Bueno Buenoseries over on YouTube. Welcome, Carlos. Got Parker from Tampa, Florida. And we’ve got India, of course. What would OpenCV be without India? And we’ve got John D from Austin. We’ve got uh Clinton says he’s coming in from New Zealand at 4 am. Thanks for making time for us in your early early morning or late late evening. Uh we appreciate it. Phil, may I may I jump in on Clinton? He’s one of our team. So uh hi Clinton, welcome. Thank you for making the time. He’s up at the craziest hours of the day. So [clears throat] wow, dedicated team you guys got over there. We’re excited to hear about what you’ve been working on. Got Picass Singh from Dusseldorf, Germany. We’ve got Mthagoram from Nevada, Texas. We’ve got uh Pamu from Sri Lanka. We’ve got uh Penny from Charlotte, North Carolina, the Queen City. We’ve got uh uh Jirona from uh close to Barcelona, they say. I I think we we are ready to uh represent there. this one representative uh from all the you know all the teams uh [laughter] we got Spain, we got Argentina, we got US, Mexico. You’re you’re Mexican now. Yeah. No, I’m team Yeah. See, Vamos Mexico. Uh team Mexico all the way over here. Um for sure. Definitely better odds than the US. Uh just from a cynical perspective. I I I don’t believe so. We can have a debate [laughter] about that. Okay. Okay. Well, I mean hopefully they will meet. That will be a hell of a game. We’ve got Muhammad from Dubai. We’ve got uh Jose from Chile. We’ve got uh Nidal from Iraq. Rene is joining us from the Philippines at 12 a.m. Thank you, Renee. We really appreciate it. And uh Clinton says, “Always a pleasure. Couldn’t give my computer or drones eyes without open CV.” Thanks so much, dude. Uh we’ve got Pakistan joining us here. Excellent. Excellent. I see there’s a bit of an issue on the Zoom, so I’ll fix that. W Sati is doing his intro. I think now is a great time to get started. You ready, boss? Yeah. Hello everybody. Welcome to OpenCV Live. At OpenCV, education is very important to us and we have our own education platform through OpenCV University. And today because of that connection, I’m very happy to invite Dennis Baldwin from drone blocks which is an educational platform for drone learning. It is uh what I call an all-in platform. So, uh we’ll go to Dennis and learn more about the platform and how it works and what kids can learn from it. But first, as always, is with me Phil Nelson who is the director of content and creative at OpenCV. Phil produces the show and if anything goes wrong, it’s his fault. Hi, Phil. Hey there, everybody. Yes, it’s me. It’s me. It’s PHIL. I am your co-host with the co-host, The Second Banana, who’s second to none. I’m also your plus one and only. You know it’s Mr. Nelson if you’re nasty, but you, my dear friends here, can call me Phil. And I’m here to remind you of a few things that we do on every single episode of this here program that you’re currently watching. The first of which is a giveaway to you out there in the audience. Stay tuned. Later on in the episode, I will be asking a trivia question based on the presentation and chat with our guest today. And the very first person to answer that question correctly in the chat wherever you’re watching. We’re live on Twitch, LinkedIn Live video, YouTube, and on Patreon will win the OpenCV University course of their choosing. You can see what courses are on offer by going to opencv.org/university.

    And while you’re there, don’t forget to check out OpenCV5, our biggest release ever, which is available right now on GitHub. I think before we get started uh with the presentation proper today um Satia can you tell the folks a little bit about how important OpenCV5 is and I’ll remind you guys we are taking questions from you in the audience during the whole show post those in the chat wherever you’re watching I think Vleta outside has a question for us [laughter] you can you can ask that question wherever you’re watching at any time and we will do our absolute best to answer those questions live during the show and we’ll save some time at the end for any of those we don’t get to during the show. But boss, tell the folks what’s so awesome about OpenCV 5. Well, OpenCV 5 is the biggest release of OpenCV in many years. I think it is seven or eight years since we released OpenCV 4. Um, and OpenCV 5 it’s it’s not a minor release. It’s a major major upgrade especially the DNN engine has been rewritten from scratch. And the reason is that when we started writing the DNN engine, Onyx standard was not a standard and we didn’t want to layer you know code after code just to support the standard. So we wrote everything from scratch. It has a new uh you know under the hood even though it will work uh similarly as the old one but it will support 80% of uh the onyx layers etc. So because of that pretty much every any pretty much any uh you know any uh model you can think of for image classification segmentation detection it is going to run on uh OpenCV transformers everything is supported right. Uh not only that it has something called a hardware acceleration layer which means that different hardware uh you could be running it on Intel CPU, it could be AMD CPU, different kinds of or or ARM and based on what configuration your uh your machine is, it will automatically pick the best optimization. So every function even things like resize uh get optimized and they will under the hood uh things will run much faster sometimes 40 50% faster than uh than previous versions [clears throat] and we are also working on the new GPU HAL which will support uh you know GPU operations all these operations will be supported through the GPU HAL so very exciting release and uh you know uh we are uh we are looking forward to getting your feedback uh on how you like it. A lot of uh things which were legacy uh have been removed. For example, we do not have the C API anymore and uh we uh you know that that makes the library slightly lighter than before and uh it also becomes easier to maintain. So please check out the library and if you have any comments, suggestions, please let us know. Yeah, thanks for that boss. And I’ll remind people so you can scan that QR code at the bottom of the screen there to go to the uh in-depth blog post about everything that’s new in OpenCV5 including a link to the migration document. There are some breaking changes there. Um you know you can’t make an omelette without breaking a few eggs and we’ve broken a few eggs with this one for the first time in a long time in fact and so it may be a little bit tough for you guys to start out but the performance gains the ease of uh writing code gains are so so very worth it. And uh we should have the uh the pip package should be updated I think imminently, right? It’s it’s coming pretty pretty fast, right, Doc? Yeah, it should be uh it should be ready very quickly. All right, that’s great to hear. And uh so before we get started with our presentation, we got one more piece of business to remind you all about. Open CV is open-source software that is produced by a nonprofit organization and as such we depend on the support of our sponsors and partners all of which are listed here on opencv.org. We’d like to take this moment to thank ARM, Futureway Technologies, Google Summer of Code, Rooflow, Orbec, the British Machine Vision Association, Jet Brains, The Edgei and Vision Alliance, Open MV, Tangram Vision, AMP Software, Tuivo, Weer.io, Intrinsic, Bears Dev, and Big Vision because every company needs a big vision. Um, I know the guy that runs that. He’s he’s pretty good at his job. [laughter] Okay, I think now is a great time to get started with our guest. Let’s bring Let’s bring Thank you, Hilda. Let’s bring our uh guest Dennis. We We had a thunderclap to to bring our guest on here. Welcome, Dennis. Tell the people who you are there. Thank you, Phil. Thank you, Satcha. I’m Dennis Baldwin from Drone Blocks and uh really excited to to be on the podcast. You know, I I was telling Satia earlier earlier, Drone Blocks is sort of centered around open-source community and we really focus on enabling uh young learners to get into uh drone technology, robotics, hardware and software. And a lot of uh the foundation that we built the business on has been uh open source and learn open CV. There’s a lot of articles that I’ve read from Satia that have enabled us as as Clinton um who’s so uh gracious gracious with his 4 a.m. time to uh mention that uh we give the ability for drones and computers to see. So, um, thank you guys for for having me. And a just a a short small shout out to Ramon from the Drone Code Foundation. He’s the one who connected us. Really excited to be here. Uh, share a little bit about our history, um, what we’re up to today and answer any questions you guys might might have. Yes, indeed. Shout out to uh, Mr. Puyo, who’s also out here in Tijuana. Mr. Pooyo. Nice. Mr. [clears throat] Poo, the chicken man. That’s his handle on uh on Twitter. Yeah, there’s a there’s actually a sign there’s here in here in Tijuana, Mr. Puyo. Um so yeah, Dennis, I think uh take it away. Tell us about what Drone Blocks has been working on. We’re really excited to hear about it. Go ahead and bring up your your screen share and we’ll uh Absolutely. under the learning tree. Absolutely. Feel free to jump in with any questions. It will start with a bit of a hardware background because what we talk about on our team is the software enablement of hardware and a lot of what makes this possible right is CPU GPU vision those type of tasks and in the early days I’m going to begin my screen share now we had the the luxury when I say we the uh developer community um drone blocks and and where we started we didn’t have to be as concerned with um the hardware platform. So this started I’d say almost uh 14 years ago just as my passion my background is in electrical engineering. I got into software and really saw uh the beauty of open source with cutting my teeth on Linux some of the uh old C++ libraries diving in and learning OpenCV but at the time [snorts] we had access to DJI hardware and we still do today. don’t you know don’t don’t rule DJI DJI I out ever. It’s just we need to be able to for our market bring this technology to the classroom. So, it started with a open-source project called drone pan, which was just a simple app built on the DJI SDK for uh the camera to be able to grab photos at different perspectives, uh angles, attitudes, and then being able to run those through a pro post-processing uh pipeline. And so at the time I was using uh DJI with drone pan and KR pano to generate some really cool uh 3D panoramas to give a a great aerial perspective. And that what we learned and I encourage all of you developers that might be um whether you’re doing stuff for the good of the community or good doing stuff to monetize it’s very easy for a neat idea an application like this to become a fe a feature in a future roll out of an OS or a software update. So, drone pan ult ultimately became obsolete because DJI introduced a click of a button panorama feature in their um DJI Fly app. So, moving forward, one of the things that was really interesting to me was mapping, you know, because at the end of the day, there’s a lot of uh landscape and construction changes that happen on a daily basis and and being able to get a drone in the air to acquire that imagery. So, I had shared on YouTube basically how to get the different altitude and aerial perspectives of that imagery and once again running it through a post-processing pipeline. At the time there was SFM, the structure from motion libraries that would allow us to stitch them, would allow us to basically create map tiles. And I’ll just uh jump in really quick to this link. There’s kind of this historic uh view of the stadium near our house. I’m just outside of Austin in Dripping Springs. And so over time, football country. You’re saying Austin or Texas in general? Friday night in general. Yeah, exactly. So you can see here that a pre-planned flight, this existed of uh consisted of about 300 to 400 images with each flight. brought them down, were able to process process them. And and just getting back to the whole theme of, you know, a drone, having sensors, having high-res cameras, being able to do something that becomes the platform, right? I I know there is a whole market out there of enthusiasts, myself included, that love to fly, that love to solder, tinker, but it becomes basically a a flying platform to acquire imagery and do something interesting with it. So that was sort of my foray into maybe the the world of open source for drones. And then moving that forward, I got involved with our um afterchool program. One of my co-founders, Marissa Vickery, she was at Walnut Springs Elementary here in in Dripping Springs. And my daughter, sorry for the ads that you know they’re going to follow you wherever wherever you go, but uh indeed, she she was running a after school program called Tech Team. And so we taught the the students how to 3D print the frame and then build their own drone. And it wasn’t long before the kids were like, “Well, let’s put a camera on this thing. What can we do with this thing other than just fly it?” So this was sort of the genesis of drone blocks, right? Being able to then take what you built and have a uh let’s say a programmatic layer. So that led us down the path of um at the time I like to I like to say that uh a drone without a camera is a toy, right? And with a camera it is a sensor. Well said. That that’s actually a theme uh one of our team members uses when we talk about we we talk to a lot of different schools and um we get compared to a lot of different hardware and we like to lean on our stuff is open source. You can build on it. you can, you know, do interesting things with it. And when we try to compete with the $50 drone from China on Amazon, we lose every time, right? So sharing that that vision and and the capabilities of what you’re you’re talking about, Satia, is very important to us. Yeah, totally. I think we lost your screen share there, Dennis. Let me bring it back. You’re back. I’m back. Okay. So, moving forward, I don’t know. Sorry, I’m not a privy to what’s being if you can see my camera as well as my screen. Sure can. Okay. Thank you. This is the T. Many of you may or may not be familiar with it, but it was an incredible platform for us to build uh programming interface, right? Visual programming is important to us because we deal with students at the youngest ages, elementary school, right? to be able to simulate something then fly a drone, take photos and um learn about the aspects of drone flight and uh programming is is is a powerful thing. So we rolled out a course using Aruko markers open CV contrib and it was met with great uh fanfare. I will say we were able to get the camera images streaming over UDP to Python and then being able to simulate or when I say simulate I should use the term emulate a lot of what DJI was to do able to do in more expensive drones the tap to fly um you know being able to do pose estimation right to use hand gestures and navigate a drone with a sub $200 piece of hardware was uh pretty incredible. Very nice. So that that shift and when I say shift we had built a lot of content knowledge understanding around uh TE platform but we had seen through many conversations with DJI where that was headed right so DJI being very uh proumer industrial military you know enterprise um hardware that that’s where a majority of their business was and we really worked hard to figure out could we perhaps take over manufacturing and as crazy as that sounds, right? But TE ultimately was sunset and at the time we had done a lot of just our own R&D around Crazy Fly. If you’re not familiar with CrazyFly, a bit craze out of Sweden, they do incredible um hardware platform. They’re very involved in the Ross community. We saw that as a opportunity for us to to pivot a little bit to make um our programming interface available to those users. And so this is our our simulator. And the the really cool thing about simulation here is we have curriculum and these different environments that students can can engage with, learn math concepts, programming concepts. But I think it’s important. I know guys, you had um I forget the gentleman’s name, but he did Pi Simverse, right? You guys had him on the podcast. And that simulation is incredible. And for us, the step that we want to take is to be able to then have parody with not only simulation, but real hardware. So that’s where we have the ability that what you’re doing here, right, with with a simulated Crazy Fly, you then could power up your Crazy Fly, connect over Bluetooth, and be able to see your your real code take flight. And I think that’s an interesting thing that I tried to reiterate as a as a developer. If you’re doing a let’s say an iOS or Android, whatever platform that you’re developing for, especially now that we have we’ll say, you know, these flying sensors, right? You want to understand that simulation does not mean 100% par with real life because of environmental conditions, lighting, wind, all of that and crazy. All that pesky entropy in real life [laughter] without a doubt. Battery’s a great example, right? It’s like, oh, here you can run it. Click connect, your battery’s dead. Right? So, all those little things um that lead to sort of the complexity, but those are teaching moments for us, right? Being able to to share that. So, moving forward, we So, this uh also has a Python interface, right? I’m looking at the blocks. Uh but it has a pure Python. It does. And unfortunately for I I don’t want to say unfortunately the thing that we have to battle with in in the classroom right everything we do is web- based right because Chromebooks and all of that so when we get to Python we’ve tried to emulate sort of fake Python in a browser right but what we really want to encourage is you know like VS code or whatever IDE so crazy bit crates has their own CF lib Python on and and we have a separate curriculum around that, but that is not um directly accessible in the browser. Okay. So, moving I love this little uh I love this little anime classroom here, by the way. This assets delightful. Yeah. [laughter] All in full, you know, transparency, a lot of what we do is we just find assets. We have a young man who’s at a university here in in in Dallas and he does a lot of uh importing and modifying the assets for us to to create this from scratch. You guys know it’s it’s it’s quite the undertaking. Yeah. Right. So I I’ll move forward to Dex C3. So we have a larger version but I’m sharing Dex C3 just because I think it’s more relevant to the conversation today. And I was telling Satia at the beginning beginning of the call. [clears throat] It it’s it’s a challenge to to to manufacture hardware, right? Especially at this level where we’re trying to solve for, you know, the educational component, right, in in classrooms, you know, we’re once again up against, hey, I can buy this toy for $50 off of Amazon or you can buy the Dexi C3, which we’re going to retail for $750. And people are like, “Whoa, you’re you’re crazy. How are we that that doesn’t make any sense, right?” And so what we reinforce is that we’re giving them a whole learning platform. And I’ll I’ll jump to that in a minute. But this drone is 3D printable. We make the STLs for all the accessories, oh, available in GitHub, right? So you can print it, you can build it, understand it, modify it so that it can, you know, suit the needs. There’s there’s the manufacturing component, there’s the build component, programming, all of that. So uh sorry, sorry to interrupt. I’m I’m because I have kids too, too. And you know, u my wife actually just signed uh my son up for some robotics thing. I unfortunately ashamed to say that. I don’t know which which one it is. uh but she signed uh so I’m very interested in uh which means that you know when she signs my son up it means that she has signed me up for [laughter] right so I’m going to be involved uh that’s why I’m interested uh like uh my question was that this uh this uh the 3D printable parts right how what percentage or what are the components that are 3D printable because uh you know when we had the uh previous drone uh which u went out of you know the DJI and uh they stopped uh they stopped producing it uh the components right the parts etc they they break frequently right it is uh it’s just the nature you have to make it light so they are not very sturdy so they break frequently uh but just to order everything from Amazon one more time it’s just a pain right so to be able to download and 3D print at home that’s uh that would be exciting But thank you for [clears throat] validating what we believe to to be true and valuable. So, uh, this to to answer your question, Satia, as you know, electronics, we’re not there yet, right? You’re not going to be able to say, “Hey, I’m going to click a button and and print a flight controller and then put it in and solder it.” But all the peripherals, I I’ll just show you the the Dexi. This is our Dexi3. We have the the top plate, the bottom plate, we have the camera mount, a battery mount, we have an LED ring. I don’t know if the viewers can see the LED animating in the background here on our larger drone, but we want to make all of that accessible so that you can replace anything at any point in time. And um we also do sell a carbon fiber version. And and that’s what I wanted to share a little bit of just if we for for the audience who are I think very technical in nature we we call it our developer kit right so we’re proponents of uh PX4 right that’s our go-to flight stack and so PX4 is where we’re running you know the the flight firmware the IMU all the attitude all that information is being handled by open source with with some light modifications from us and then we add basically what we call our our code stack. So this ships with a CM5. We’ve designed our own carrier board for breakout, you know, for the LED or any GPIO that you might want. And then we have a a bunch of content and I I’ll I’ll jump into that here in a minute related to doing vision tasks. We we started as I mentioned with the TE and Aruko markers and more recently we’ve been shifting to doing um for the audience solving the problem of of drone positioning outdoors I’d say is significantly easier right than solving it for indoors. Outdoors you you plug in your GPS receiver. Now I’m not talking about drone light show. We’ve all seen that. That’s that technology is amazing. high-end GPS and accuracy, but a standard GPS, you have the ability to just get your position, right, 3D position, and then program your drone. So indoors we make use of optical flow and a distance sensor, but really the higher fidelity or higher accuracy positioning in my mind has to come from something like fiducials either like April tags or ruko markers you know um around the environment or Satia I know you’re you’re you’re a slam guy right so we’re we’re look looking at doing some pre preliminary vio with slam Right. Moninocular slam with a single camera would be amazing. Yeah. But I could ramble on and on about the complexities, but the thing that we want to do is sort of help solve for, you know, giving students the ability to build these technologies, understand them, right, and then jump into any of the repos, right, that we provide, our April tag, our YOLO, all of that. We make it available. We bundle that into just a we call Dexios. Dexi is our drone for exploration and innovation. But that that bundle th those packages are just, you know, Raspberry Pi OS with Ross 2 and all of the different libraries and packages built into it. Quick question. Have uh do any of the drones have speakers on them? And have you made any of them play Come on Eileen by Dexis Midnight Runners? [laughter] So, uh, no, but here’s what I will say with the carrier. I’m going to, um, just hold this up to the screen, Phil. I hope that the audience can see this. We have some breakout ports, right, that sit on the sides here. Sorry if that’s not visible. Yeah, you can see it. You would have the ability to whether it’s I squared C or whatever the interface of that speaker. The the the problem with the speaker is that the drone is also fairly loud, right? So, yeah, that that that great sound. Yeah, exactly. [snorts] Okay. So, I’m going to share this last [snorts] um piece that is I I would say near and dear to our team. It’s called AVR, Advanced Vertical Robotics. So when you come into the drone blocks ecosystem, you have the ability to go through the the different tiers, but one of them is a drone competition that started out of Bell Helicopter and in the DFW Fort Worth area, but now they’re known as Del Bell Flight, right? You you guys might be familiar with the Osprey. It’s a really cool aircraft, vertical takeoff and landing. Bell gener uh initiated this competition eight years ago and we got involved about four or five where we started bringing our own uh hardware to to the to the competition. So you have the ability in the platform to spin up the simulator and I’m just going to show a little example. So this is the simulator on the left and this is the 2025 game field, right? And so the student that keep in mind just for for the audience’s benefit that all of this existed physically before were teams from all over the country predominantly USA would compete in regional events and then the top let’s say 15 end up in Dallas and there’s a finals where they’re learning how to build solder program and and solve uh various challenges for time if you will Right. And so what is the team size? Usually it’s it the cool thing about it Satia is like so I don’t you mentioned your your your son or your daughter doing the the camp. Yeah. Those generally, you know, are probably first robotics. That’s my guess, right? FRC. And so what Ron Olsen, he’s the um guy that created the competition and just has a very uh creative mindset was like, “Hey, this competition needs to be different by including more uh players of the team. So a team should be a minimum of eight students, right? Because there’s a ground robotics part where they’re navigating a maze. um there’s somebody who’s sort of overseeing the team captain and then there’s multiple pilots. In this scenario, there’s just the one dexi, but in the real world competition, there’s a large dexi and and this smaller one and the kids the competition is by the age or how does it it it’s generally targeted at high school. But what we’ve tried to do with the simulation is expose it to a younger audience, right? because you you know in simulation for example I’ll just you know navigate around and I’m going to show you guys that we have a few April tags spread out through throughout the course there’s some images so it just scan the April tag and I I have code running in the background this is all node red on the on the right that that layers into a ROSS 2 kind of uh foundation [snorts] but those you know students we want to make it accessible at a younger age so that they just get the experience of learning how to fly, learning the different concepts and you know you don’t have the maybe the overhead of uh understanding all the electronics initially right if if the competition and the programming sort of piqus your interest I believe that just my own kind of mindset is as an engineer if I if I have some level of success at a higher level write some code and it does something then I dive in a little bit deeper if that makes sense. Right. Right. So, I’m going to just jump over really quickly. So, we have these inside these cabins. There’s a whole story behind this game. There’s stranded animals and then you look in here and you sort of scan for the animals and that’s all doing yolo um detection. So, we have this, you know, Onyx model that we’ve uh did some transfer learning on. Guess I’m I’m probably not doing a good job of flying here, but okay. Yeah. So there there’s the bird, right? And so gosh, apologize to you guys. Target scanned. And so what happens is we have all these ROSS nodes running. And I want to I want to wrap with this really quickly. And one of them is um a YOLO detection node. So we have our April tags and then YOLO detections. So, we’re teaching, you know, students, right, to to be able to to scan, to wire up their logic, right, to determine the confidence level. And then they want to go deeper in the future. It’s like, hey, we’re going to put completely new objects out there. You need to use whatever model you decide to use. uh fine-tune it to to detect some new or add new objects to the to the model to be able to train.

    Very nice documentation. Nice. [snorts]

    All right. So, yeah. Um we got some good questions here, but I think and you you you your timing is is just about right. We’re at 37 minutes here. So, not not bad at all. We got quite a few questions here. We’ll get to Sorry if there’s a little bit of background noise. Um, I’m going to quickly do our trivia giveaway here. Um, a little bit early in the show. So, solo myself. So, if you’re joining us for the first time or you just need a little reminder, one of the things that we do on every episode of Open CV is a giveaway to you in the audience. I am about to ask a trivia question based on today’s presentation and a little bit of the stuff we talked about. And the first person to answer that question in the chat wherever you’re watching on Twitch, LinkedIn, YouTube, Patreon, or on Zoom will win a free OpenCV university course. These are awesome industrial strength courses created by people from the industry to help you improve your understanding uh level up your career or uh get a better job, get a new career. Um so check it out at opencv.org/university.

    And here comes our trivia question. If you have won in the last couple of months, please don’t answer. Give someone else a chance to win that free course. I’m looking at you, Martha.

    So, in what year was OpenCV 4.0.0 released? We talked about this at the beginning of the show that it’s been quite a few years since major releases of OpenCV. In what year was OpenCV4 first released? Answer in the chat now.

    Oh, wow. Axel Tracy just boom out of the gate. You get out of my head, Axel. Uh the answer we were looking for there was 2018. The answer was 2018 2018. So it has been quite some time. Uh congrats Axel. Were you just guessing? [laughter] He says in in chat. Wow. He says yes. I just guessed. [laughter] Hey, it still counts. Um uh congratulations. We will uh so please go to opencb.org/university and choose one of those courses. Then send exactly one email to me, philopencv.org. That’s philopencv.org. And we will make sure that you get that course unlocked in your OpenCV University dashboard. Uh, congrats, Axel. He says he uses he used context clues. Satia said seven or eight years ago, and there we are. Yeah. So, congrats. Um, and I think let me go ahead and remove our thing here. Okay. So, we’ve got uh quite a few questions here. Let me pull up pull up the best one. A couple pe couple people were close. You you’re lost by maybe three five seconds. Uh cat 5Dev on YouTube. Um Penny on YouTube and Emmanuel on uh LinkedIn. Uh better luck next time, guys. Um so questions. Our first one here, uh, we’ve got, uh, YouTube asks, “Can you connect to the drone from the web, uh, or do you need to install an app on a smartphone?” Dennis, that that is a I wish there was a simple answer. So, I will say that if you’re flying with a TE, uh, definitely that was native because we had to use a Wi-Fi bridge. So that would have been an iPad, an iOS, or even our Android app for Chromebook for Crazy Fly. We use uh the Bluetooth web API, and you’re able to connect to that drone uh via web browser. And then now here we are with our own hardware with Dexi. When you boot Dexios, given the fact that it runs on a Raspberry Pi, what happens is that we bring up all the ROS services, we bring up the April tag, the YOLO detections node, we broadcast a hotspot, you connect to the drone over the hotspot or you put the drone on your network and all of that runs uh in your web browser or on your mobile device. We we aim to make our web interfaces as responsive as possible. what what we’ve struggled with I think the the the difference is you know being a developer doing sort of commercial stuff before drone blocks was oh we’ll just make you know native apps right and generally that becomes a bigger challenge at least in the US education market to allow third-party apps to be installed um on student devices and so that’s where if we can make our technology web accessible. Uh that that wins every time. Yeah, they tend to keep those Chromebooks locked down pretty hard for, you know, for obvious reasons. I It makes a lot of sense. Uh kids are going to kid. They’re going to they’re going to figure out ways around whatever blocks you put in front of them. I I’ll tell I’ll tell Phil, I got to share this. Please do. Side story. It’s It’s not going to be 100% accurate, but it’ll Ryan from our team supported this use case. So we have a native I’ll say a chrome app right so you used to be able to develop chrome apps that you install from the chrome web store uh chromebooks and so we used the standard uh Google authentication right to log in and some way shape or form during the on the off screen there was like a privacy policy or whatever and kids figured out how to click that took you to a page and all of a sudden you got to like unfettered use of the unrestricted use of the web and some kid made a Tik Tok video and then all of a sudden we saw inst installs spike and we were kind of like well maybe we’re in the wrong market right maybe we need to help kids figure out [laughter] how to get online so it was I love that pretty wild life finds a way you know totally that’s that’s what I love about our audience though even the the drone AVR sim that I was showing you kids are finding ways to to find loopholes Right. And I think that that ingenuity, that creativity, creativity is necessary. So, yeah, definitely. I mean, I think about when like back to when I was a kid, you know, we were figuring out how to install like uh, you know, Doom and uh, Sim City and stuff on the school PCs and and how to like, you know, set up the toggle so you could just like alt tab quickly over to the, you know, the Microsoft Word doc to make it look like you were actually doing your assignment and stuff, right? And I think about that like, you know, back then sometimes the teachers didn’t, you know, we knew more than the teachers. And I think that’s that’s definitely still the case, you know, today, especially today because technology moves so much faster now that the kids just instinctively know this. And so if you’re trying to, you know, any constraints you put in front of them, they’re going to figure out a way around it. And I love that. I think it’s a good sign like uh just the yeah as you said ingenuity and uh you know it’s it’s something that we reward in adults and sometimes unfortunately punish in children. Yeah. And uh yeah, they they find ways like uh my my wife uh you know uh closes down the Wi-Fi for my kids laptops, right? Uh to restrict time and uh even the younger one, he uh thought about it, how to bypass this problem. And so he swapped the names of his MacBook Pro like when he had the access to her phone, he swapped the names of his MacBook Pro and uh you know some some random thing like the water heater or something like that which is also connected to right and it shows up as the water heater, right? Swap out the name. So when she pauses the water heater, it’s paused. He’s still on his [laughter] right. That’s fantastic. So his MacBook Pro, the name is swapped. So, uh, they do all these, uh, things to, you know, I love that. If you’ve got any similar stories in the chat, please, uh, drop them in there. We’ll bring them up on screen because I I love that kind of stuff. Um, we got some more questions here. Um, let me go ahead and bring up the chat again. Um so uh Jurgen Fay asks I think a really good question here and especially this is I think a thing in you know in the United States where education is is vastly underfunded especially public education. Um Jurgen asks how do school folks deal with the cost even five or 10 per class are an investment for most. What would be the best age for starters? So I think these are just two questions but these are really good questions. Yeah. Well, h thank you th those are those those the first one makes my my head hurt but and I and part of it is because being a co-founder engineer evangelist whenever I’m faced with the question of how much it costs I like defer to someone like Lindsay or somebody else on our team sales is not my forte but I will share a little bit of what we’ve learned over the years at least here in the state so one there’s there’s Perkins grants right? There’s a Perkins grant, you know, from the federal government and that gets allocated to to schools that apply. So that that’s one, you know, grant funding that exists and we’ve actually worked in the past to sort of say, “Hey, school, we we will help you um with the grant application process. Marissa and our team has experience writing grants.” So, so that’s one thing. The um other thing is we’re part of this u I don’t know if it’s movement shift in education to this whole CTE uh space where it’s career and technical education maybe you know less students who are graduating or going off to university but they’re actually finding their calling maybe sooner or maybe they don’t know what they want to do but they’re not ready to invest you know what what a a college college college education cost nowadays. So, the career and technical education in computer science, there’s a lot of um funding and programs available, grants available in that market. And the other thing, the the final point I will make just based on what I know, there is a sweet spot for a purchase that’s maybe right under the the threshold of $5,000, right? Which is still a lot of money. I I understand that, right? But if you break things down, that’s sort of where we’re trying to to package our technology and and software. And and the thing that I will reiterate and the thing that I believe is is is tough for us to overcome is that we’re constantly being compared. Once again, getting getting back to someone’s gonna go put a programmable drone in their Amazon shopping cart that’s like less than $200, right? And so if you’re like, “Hey, I can get five of those for a,000 versus five Dexi3s for 4,000, it if it’s just a numbers or a math game, we’re going to lose every time.” So that’s where we have to reinforce just like you guys do with the platform and the education. uh the the soft skills, the the learning that you’re going to have access to when you don’t have hardware on your desk and in your hands. But um great question. And the the second one, what would be the best age for starters? Um if if I may, Phil, may I may I share again? Absolutely. Okay. So, part of what how we’re positioning this now is we have ground school, right? And and that’s where we’re focusing on like the the K through five, right? So you have we have our flight academy which is just a you know a simple hey here’s here’s how you might interface with with Dexi. Um well this is how you interface with Dexi before you actually have one, right? So you have all these challenges, right? You learn how to arm and disarm fly. So we try to make that accessible in the browser before they get to real hardware. So, we have a bunch of content and challenges around um simulated drones being able to do different uh tasks. This flying curves course is is incredible. This is from a advocate in um Italy, Luigi. He’s always sending us really cool content, right? Stuff that you can do with these lessons. And so, you’re learning, you know, different concepts of programming as well as math and you’re able to apply that in the simulation. So, our hope is that we make this accessible at at a at an age, ground school, K through five, that piqus their interest. So, when they get to middle school, getting their hands on the real hardware um becomes a lot uh more accessible and easier to understand.

    Yeah. Um great great answer there. Thanks for that context. Um more questions we’ve got. So, let’s’ll take a break from questions here for one second. Um, uh, Satia, I want to give the folks a little can we give them a little sneak peek of the pretty big announcement we’re we’re going to be doing in the next week or so um about the next OpenCV competition. Uh, without giving it too much away, I think we should uh clue them in a little bit. Yeah. So we are this close to signing a contract with uh one of the big u you know um big uh hyperscalers let’s call them um and we are working with them to create uh for a competition where you will build um applications just like our previous competitions we we will uh ask you to build applications and it could be in two stages. We are still deciding on the exact format of the competition but uh typically we do uh competitions in two stages. In the first stage we ask you to uh to basically propose a solution or propose an idea for a problem that you want to solve and based on that we select a few and then we will give them some credits on this hyperscaler and based on that you will build the final solution. So it’s uh it’s going to be interesting because you know a lot of folks would be using OpenCV5 in the process uh and we love that because you know it’s getting so much faster and uh we have divided the competition that we’ll we’ll let you know there are a few use cases that we are very interested in right if you want to use it um as an agent uh that would be that would be great if you create an MCP server that kind of thing uh it would uh there will be a few categories uh where you can compete Those are suggestions and you can pretty much compete in any category. Right? So a new competition is coming. It would be uh based on one of the hyperscalers and we will uh basically make some credits available to you for people who qualify the first round. So as exciting as always uh we’ll you know uh it’s it’s going to be fun

    indeed and I’m really looking forward to this one. Uh, as Satia said, we’re pretty close to being able to formally announce it and it’s going to be uh a little bit different than some of the previous competitions because of the uh the partner we have this time and it’s I think maybe more accessible than ever for for people to participate especially in countries where it’s hard to get the hardware or it’s difficult to access compute and and stuff like that. So, uh we’re really looking forward to it. Um, I’m also going to take this moment to remind folks that uh OpenCV has our official merch store um at opencv.mmyspreadshop.com and we just released a brand new t-shirt which I am going to bring the QR code up on screen for you guys. I think you’re going to like this one. It’s uh near and dear to our hearts. I will cover up Dennis’s face briefly, but you can buy official OpenCV gear on the OpenCV merch shop at opencv.mmyspreadshop.com. We’ve got a brand new shirt there. Says open- source, nonprofit, no [  ] That is the OpenCV credo. Uh we’re we’re leaning into it. We can uh get away with things that maybe some of the larger organizations out there cannot. And uh that’s part of the power of open source nonprofit software is that Phil’s allowed to say [  ] on official merchandise and on our live stream. [laughter] You you said worse things in in videos. That’s right. I’ve said worse things to you about you. [laughter] But please scan please scan that QR code and check it out. We’ve got hats, t-shirts, the OpenCV logo on all of that good stuff. We’ve got little bandanas that you can adorn your pets or small children with, as well as stickers and pins and all that good stuff. So, please check out the official OpenCV merch store. And if you buy something from the shop, take a picture of yourself wearing it. Take a picture of wherever you stick that OpenCV sticker and we’ll share it here on OpenCV Live with your permission, of course. Okay. And we’ve got a few more questions here and a little bit of time to answer them. Let me pull open the chat once more. Uh, we’ve got a question from Clinton Evans, who I think knows the answer, but uh, we’ll we’ll allow this uh, putting the thumb on the scale a little bit about how long should a drone kit last in the classroom on average, Dennis, this is a good question for those educators out there. Uh, a drone from drone blocks. Yes, I would say from others. We we want to we want to create longevity. So, uh, our our kind of phrase or tagline is elevating K through life, right? Get bringing them on the continuum. And so, in our scenario, if you’re introduced to the Dexi3, we we we bring that in around middle school, that should electronics wise, battery, you know, uh, motors wise, whatever. I don’t really it it’s hard to put a shelf life. I would say 3 to 5 years maybe on the electronics batteries depending on use, right? But the thing that I want to reinforce is that you have the ability to 3D print as we discussed earlier and add modifications. So we we’ve had some really cool contributions from the community on like claws and grabbers that you can attach uh bring up a servo, right? you know, wire it up with Node Red, Python, whatever your your interface of choosing and extend it, right? So, I would say, you know, a three-year time frame is what I would feel comfortable just, you know, off the cuff saying from from an electronic standpoint, but the the repeatability of printing and modifying, you know, can exist uh indefinitely. Yeah, that’s one of the really nice things I think about the platform you’ve built here is I mean you basically can you know if a servo burns out put a new one on there right that’s uh that’s very nice [snorts] what is the difference between a developer kit and the single build kit good question Satia like one of the things that we have been doing I I have a I’ll just call it the the flight kit it it’s just our our flight controller with optical flow and so there’s There’s no true programming, if you will. It’s learning how to fly, learning your different flight modes, navigating around. And then the the developer kit is where we introduced um the the Raspberry Pi CM5. Uh the carrier for more breakout and expansibility. One of the things that we do have uh for a larger drone and we’re trying to bring this down to a smaller scale is the ability to um add in this case we have a hat with a this Halo accelerator, right? So we we we right today that most of the YOLO that we do and vision stuff runs on CPU and that’s as you guys know constrained for a Raspberry Pi and we can slow we’ll slow the frame rate down but with acceleration we can um you know obviously run more uh frames for inference and so really the the flight kit is sort of the standard you you have to learn how to build and fly Right. So, if you that’s that’s what we reinforce and then when you get into the programming um that’s where you have the the add-ons with the compute and accelerator. Got it. Got it. Very nice. Um got another question here from Zoom. Um Zoom asks, “What’s the possibility of having a few different cameras facing down on one of these drones?” That’s a I wish I could untether from my headphones, but we have actually I might be able to do this in real time to um I I I understand the question. A few cameras. I’m going to share my screen if that’s okay. Phil, can I do that? All right. Here we go. And so, um this is for the downward. Let me just say down. Okay. Yeah. So we have a downward camera mount that that sort of mounts in the frame. We right now run py cam V2 and V3. And so you [snorts] there are two ports two camera ports on the I’ll say DEXC5 that’s the larger drone 5- in propellers that you can basically uh run through the the camera stack. Right. So you can do t two simultaneous cameras with the DEX3. Right now, we just support one and with with a tilt modification, right? So, you have the ability for the forward facing or the downward. And we’re working to uh have multicam support for the smaller drone, but it doesn’t exist today. Got it. Yeah, thanks for that. Um, I’ve actually got a question from me, which is I’m always really interested in the thermal issues for this this kind of embedded compute because it’s always an issue and and even with like the Raspberry Pi, especially with the Pi 5 uh and and the the CM5, they get a lot hotter than previous way hotter than previous versions. And so, uh, what I’d like to know is do you notice a like what are the differences in the thermal temp with say a drone at rest versus a drone moving around up in the air? Uh, if if you would have told me that question were coming, I could show I I would I’d spend the rest of this call trying to trying [laughter] to find that chart, right? But so CM5 is crazy hot, right? Right. And so we’re we advocate if you’re doing development, one of the things we do is advocate do your development in the simulation, right? Like that we have code VS code in the cloud, whatever. When you’re ready, bring that over to the drone. But I’m guilty of just testing on real hardware sitting on the bench. So we have like, you know, a cheap USBC CPU fan or GPU fan that just keeps it cool. Now, I would say under rest, we’re probably at 80 to 90C, right? And and we’re not really taxing the CPU very hard uh with our out of the box OS. But what you’ll observe is that the minute you arm the motors and because of all that, you know, propw wash we call it, as you get up, we’re we’re we’re down to 50C or less, right? While it that’s a huge drop. And I would there’s ever an opportunity to to share in the future. I I’ll I’ll send you a link just because you asked, you’ll see that just dip so quickly. Much faster than, you know, obviously a CPU fan sitting on the bench, right? It’s um Yeah, great great question. Oh, thanks. We’re we’re I think about at time here. Um, if we’ve got any last minute questions, let me go ahead and check out Zoom. Um, looks like I think we got to everything here. Um, thanks so much for joining us this week, Dennis. This was a lot of fun. I I’m really uh you you all have built a what looks like to be a a really thoughtful and uh well-made platform for for education and for especially I’m I love that what you’ve the thought you’ve put into you know the repairability and the the DIYness of it all because that’s also part of the education process like there’s something really exciting about uh oh you know one of the propellers broke well okay just go 3D print another one right like figure out the slicer and go 3D print another one like I I love that. Um, Phil, I know we’re running on time. I’ll just No, it’s all good, man. We We can The dream would be to 3D print and not have to injection mold propellers. Um, we’re still at the place from efficiency standpoint. I know you were using that example I’m getting hung up on. We would love to 3D print propellers that that were very efficient, didn’t kill battery life, and um but uh yes, everything but I’ll say that. Gotcha. Gotcha. So, the one example I chose was the wrong [laughter] one. But guess what? We have a gentleman that does a lot of CAD work and I’m going to tell his name is Aziz. Does incredible. He designed all our frame and everything. I’m going to say, “Hey, the Open CV Live guys uh said you couldn’t design and print a efficient propeller.” And we’ll see what he says. Oh yeah, sure. Blame me. I um Satia, we’ve got an awesome episode coming up next week with our friend France. Um I hopefully he’ll be wearing his raspberry beret. There’s another little musical reference for you guys out there. Um do you want to talk about cloud optimized OpenCV which will be our topic next week? Yeah. So uh next week we are going to talk about you know cloud optimized OpenCV. We some of you may have already heard about it. This is our uh optimized version of OpenCV on AWS Graviton. So uh Graviton’s new version Graviton 4 is out and so we are going to see how OpenCV performs on these processors. A ton of other things that France uh will is more more uh capable of uh explaining. So we’ll have that episode next time. So tune in. Uh it’s it’s going to be great. Yeah. Scan that QR code I just put up on screen. That’ll take you to the Amazon Web Services Marketplace listing for cloud optimized OpenCV. And you can join us next week to learn all about that. Um, I’m also going to remind you OpenCV has a Patreon. You can join us for just $2 a month and get OpenCV live streams over on Patreon as well as DRM free downloadable episodes of every new episode of OpenCV Live. Watch them however and whenever you want for just two bucks a month. There’s also a $7 a month tier. We appreciate all of your support. Everybody that gives us a dollar, $5. It really means a lot to us. Everybody working at OpenCV, not just myself, not just Satia, but the whole team out there. And uh we’ll be talking with a couple more of them. You guys have gotten pretty familiar with Gersomer and Abashek. and I I think they’ll be uh possibly uh be joining us a little bit next week as well to talk about cloud optimized OpenCV. But scan that QR code, join us on Patreon. You can sign up on the Patreon for free, but you won’t get those downloadables. You will, however, get um all of the Patreon posts and stuff like that as part of your uh free subscription on Patreon. We’re going to be trying to leverage this a little bit more. And hopefully in the future, we’ll also be now that Apple podcasts has video podcast support, we’re looking into uh setting up uh video podcast on Apple Podcast as well. That’ll probably be a paid thing, but it it won’t be expensive, I promise. So, please join us on Patreon and uh of course, subscribe to the OpenCV newsletter. Um follow us on the various social media platforms. We’re on Masttodon, Blue Sky, LinkedIn, uh YouTube, uh pretty much everywhere. And uh last but not least, thank you all for joining us this week and thank you Dennis for the presentation. Take care of yourselves out there, folks. And take care of somebody else if you can. And vamos Mexico. Adios.

  • 06/18/2026 – OpenCV Live

    Hey there everybody. I think we’re live now. I do believe we’re live on the wild and woolly world of Al Gore’s internet. It’s Thursday morning. It’s 9:00 a.m. and you know what that means. It is time for OpenCV Live. We’ve got a few folks popping in already over on YouTube. Love to see it, folks. And uh we’re live on Twitch, LinkedIn, YouTube, and Zoom. I’m here with uh the homie from Rooflow who’s I think this is is this your second or third episode of the show? I I think second. I think second. Okay. Well, welcome back. We’re pleased to have you to talk about this awesome bit of technology from longtime OpenCV supporters. Rooflow RF DTER has been a huge huge thing over the last you know year especially and uh lots of people were talking about it at CBPR this year. Um, I’m like, “Hey, I know those guys. They do they do good stuff.” And so, we decided to have him on the show here to to do a little chitchat. Please, as you join us, wherever you’re watching us, let us know where you’re joining us from. I am coming to you from beautiful Tijuana, Mexico here south of the border. And it’s uh looking to be a pretty dang nice day outside. Maybe we’ll take the dogs for a walk today. Um, looking forward to that. But let us know where you’re joining. Looks like we got Nevada Tahas here. Mathagorum chiming in. Uh we’ve got Jurgen on YouTube says, “Cheers. Awesome work since years.” We got uh and Jurgen’s coming from Munich. Yeah. Right on. It’s an international affair here on OpenCV Live as per our usual. Um, so over on the Zoom chat, we’ve got uh Stefan from Vancouver, Canada. We’ve got uh Milos from Serbia. We’ve got Zack Crane admitting to be coming to us from Iowa, which is a brave move on on Zach’s part. I’m I’m We appreciate the fact that you um stumbled out of the cornfield to join the show today. We got Allan from Austin, Texas. Love to see it.

    We’ll get started in just a minute here as we let people uh get into their viewing situation. And as you can probably hear one of the dogs yaking on something in the background there. Sorry about that. Poor uh Fluffy. She’s on some meds now, which are helping her out. We got uh Portugal. We’ve got Slovakia. We’ve got India. We’ve got Kiev. We’ve got the other Vancouver in Washington. Lots of folks in today. Yeah. Got Hungary. We’ve got Picash from California.

    Uh we’ve got uh somebody joining on a train in Germany on the way home. Awesome. Right on. Thanks for making time for us in your commute, uh, Valentin. We appreciate it. We got India as well. Uh, somebody says, “I love you.” Oh, that’s very that’s very kind.

    Thanks so much everybody for uh participating in our little where are you joining from segment this morning. It’s awesome to see all of you. It’s one of the best things about this show. It’s the Open CV audience and community are so freaking cool. Looks like we’ve got Caro from Pittsburgh and then he’s uh but he’s based in Austin. We’ve got uh Andre from Moscow uh in uh Kimiki Kim Kim Kim Mickey I think. Uh we’ve got Molly. We’ve got St. Louis, Missouri. We’ve got uh Riad, we’ve got Chicago, we’ve got Birmingham, and we’ve got Olaf from Costa Rica. Thank you guys.

    Going to slightly tweak my microphone here. I think uh you’re getting a little too much background noise. Nobody needs to know that much about what’s happening here.

    All right, I turn I turn turn it back on here. Hopefully, it should be better now. All right, so um I think now is a a great time to get started. Let me go ahead and solo myself and welcome everyone to the show.

    Hey there. Hi there. Ho there everybody and welcome to Open CV Live. It’s Thursday 9:00 a.m. and it’s time for our show. We’ve got a great one for you today. We’re going to be talking about RF DTER key points, the new amazing preview release from longtime OpenCV supporters, Rooflow. This one is really exciting. It is proving to be uh potentially even more significantly more accurate than YOLO for various tracking instances. The one they’ve got now is is key points. And uh we’ve got our friend from Rooflow to come talk about that with us today. But before we get started, I’ve got a few things that I need to talk to you about. The first of which is a reminder that OpenCV is a nonprofit organization that releases open-source software. And as such, we depend on the support of our members and sponsors, all of which are on the screen right now. We want to give a big thank you and shout out to ARM, Futureway, Google Summer of Code, Rooflow, Orbit, the British Machine Vision Association, Jet Brains, Intuitivo, The Edge AI and Vision Alliance, Open MV, Tangram Vision, Amped Software, Intuitivo, Rerun, Intrinsic, Beeris Dev, and Big Vision because every company needs a big vision. If you want to be as cool as these companies are, the best way to do that is to scan the QR code at the bottom of your screen and sponsor Open CV, you can Where’ the Dang it, where’d it go? There it is. Scan the QR code at the bottom of the screen and sponsor OpenCV. You can donate as an individual. Or if you work for a large organization, it’s very possible that your boss can turn your donation to OpenCV into a 2x donation. For example, if you donate a h 100red bucks and your company participates in donation matching through a system say like uh Benevity, which OpenCV is a part of, that 100 bucks can turn into 200 bucks in the blink of an eye just by talking to your boss about it. We hope you’ll do that. But even if you can’t do that, there’s other ways you can support OpenCV, such as by buying OpenCV merchandise. We’ve got official t-shirts, hats, tote bags, etc., and bandanas for all of your pets and or children. You can also sponsor us on GitHub at uh github.com/sponsors/opencv.

    Highly encourage you to do that. It is a great way to show your support and get a little bit of uh uh social capital out of having your name on the OpenCV supporters page, which I will read out a little bit later in the episode. So, please scan that QR code, help out OpenCV however you can. As I said, we’re a nonprofit that makes open source software, and there are not a lot of us left, especially in the computer vision industry where a lot of companies are even taking what was open- source and making it uh closed source or restricting the source to for various purposes. Uh we don’t want to do that. We’re not going to do that. We are here for you. We are by the people, of the people, and for the people. And we appreciate every little bit of your support. So, we’re also taking questions from you in the audience. Please use the Q&A button if you’re watching on Zoom to ask your question at any time or just post it in the chat if you’re watching us on YouTube, LinkedIn, Twitch, etc. I’ll be monitoring those chats and I’ll bring up those questions as the show progresses. We’ll also save a little bit of time at the back end there to answer anything we didn’t get to. So, please do that. We’ve also, as always, got our trivia giveaway later on in the episode. I will be asking a trivia question based on today’s presentation and the first person to answer that question correctly will win the OpenCV University course of their choosing. You can see what courses are on offer by going to opencv.org/un university. We hope you will. It’s a great service and uh we have a lot of success stories on there for you to check out. Maybe you’ll be the next OpenCV university success story. So stay tuned for trivia and check out OpenCVU. That is about enough of my yappen today. I think a couple more folks chiming in. We’ve got uh hello from Liberia. We’ve got Bavaria. Um lots of the IA countries chiming in here. We’ve got uh Germany, Poland, got Austin, Tahas, and we’ve got uh Wesley from Oregon. Notice I I pronounced it correctly. Wesley Oregon. You’re welcome, buddy. All right. I think now is a great time to get started. My friend, introduce yourself. Remind everybody who you are and tell them what you’re going to be talking about today. Uh, hi everyone. My name is, uh, Peter. I work at Rublow. Um, and more or less all I do is open source. Um, and yeah, today I’ll be talking to you about our PTR and specifically about the new release that we made. Uh we added key points this week. So that’s the main topic.

    Indeed. Uh you can go ahead and share your screen using the button at the bottom of the screen there whenever you want and I’ll pop it up. Let me do that. Let me do that. Share screen.

    Dar she blows. If you’ve got audio on here, I recommend turning it. We don’t have audio, but we have some videos. Uh, smart man. Yeah, finger crossed everything will work. Um, cool. So, I guess I’ll just proceed. Um, yeah. So um last year uh we’ve released uh RVtr it’s uh it was but it’s no longer is uh only a object detector uh right now we support more tasks and that’s why we meet today here. So uh RFDTR as u the object detection part is state-of-the-art object detector uh beating uh other top choices for object detection uh both in terms of speed and accuracy. Um so here’s the benchmark on the cocoa data set where you would like to be here is in the top left corner that would mean that you are highly accurate and fast model uh for object detection. Um yeah, that’s uh what we aim for. But not only that, um another kind of like a benefit of RFDTR is uh let’s go NX. Yeah, I agree. Um uh another benefit of RDTR is that we are very good at uh fine-tuning. Um so here’s another benchmark that we internally created. Those of you who know Rooflow, you probably know that we have a lot of data sets on the platform. Those data sets are uh created and shared by our users. Um so what we did is we picked 20 of those data sets. We uh took a look at you know uh what’s inside and and um um collected sorry not 20 but 100 data sets divided them uh into into buckets and um decided to check how well different uh detectors can fine-tune of on those data sets. Uh and it turned out um is the top in that category as well. Uh so uh here is kind of like the average score. Um but if we took a look at uh individual um buckets uh we can kind of like dive a little bit deeper. So on average you can see that for example if you would choose uh YOLO 26 um and RDTR at the same uh speed you can on average gain around 2 map. Uh so nothing else changes. All you did is you swap model from one to another and during the inference uh you can you know kind of like for free get um this accuracy boost but depending on on the category on the of the data set you can uh get even more. One of my favorite uh categories is aerial. Um and on aerial data sets, RDTR uh scores uh very often around even like five map points higher um than uh you know those other object detectors. The reason for that is um RDTR uses Dino V2 backbone. Um so it uses uh also open source model from Facebook as a backbone. That model was trained on like crazy crazy amount of images uh very diverse um images and that helps the model to learn fast because it already saw a lot of those things in the past um where contrary to uh you know other uh popular object detectors that are pre-trained ImageNet or maybe on Koko all they all they saw is like a very like a very very narrow u part of our life and dino saw uh aerial images so medical images so a lot of other things so um that knowledge is already in RFDTR because it’s already in Dino V2 and during training we just we just you know get that information that is already there um so that’s that’s an example of what you can do with uh RDTR with like aerial images uh or aerial videos Um, and the model trains a lot faster uh than other uh open source models. Um, so it it it is slower per epoch, but you need I don’t know like five epochs to already get like a very very good accuracy that you would need to take probably like 25 or even sometimes 50 epochs uh to get with other open source models. Uh so that’s a significant amount of compute to save. I mean that’s that’s no joke. Yeah, exactly. So So um this is like one of the main I would say benefits of the detector that that we released is that it’s really um really good at like quickly getting to proper results. Um, and you can see like if you compare like yellow models uh and RFDTR when you train, it’s usually like RFDTR is like a steep wall during the first five epochs and getting to like very very reasonable uh accuracy and it takes YOLO a lot of time um to get there. Another uh interesting thing is that uh you can get those results with without almost any augmentations. So those of you who are familiar with um how we did object detection for the past uh several years, many of those training grants, it was just like a creative way to augment images um to introduce enough variance in your data set that the model um can learn and uh generalize later on. uh RDTR does not use augmentations at all. Uh and still can learn pretty fast because like I said that knowledge is already there. All we do is we just look for you know where like try to get to that information that was already in that backbone. Um and yeah uh then few months later we released the segmentation uh model and the segmentation model is uh once again uh state-of-the-art when it comes to speed and accuracy especially when applied to uh sporting events from ESPN 8 OO. Yeah, I mean uh that’s you know one of the ways that they like to play with those models uh for sure is is applying that to sports. Um yeah, that that one is very interesting the uh word chase tag. I believe that was professional tag, right? Wasn’t that guys who who participate? Yeah, those are uh like parkour uh right and then they are doing crazy things. Yeah. So it’s actually pretty hard you know to track them because they are move in a very unpredictable way but that’s a that’s a separate topic. Um so yeah so yeah we released segmentation mode uh once again getting um getting state-of-the-art accuracy and yeah a lot of people um asked us if we plan to release key points model and uh this week we released the preview version of the key points model that’s what we did with segmentation model in the past so instead of just going all in and get giving people u all sorts of ways. What we do is we we prefer to give people like a single checkpoint, ask them how do they use that, you know, uh what can we do better and then release the actual um like a proper proper release few weeks later. So that’s definitely coming. We intend to give people like a full uh range of uh checkpoints. I think four or five uh different sizes, but for now we released the the largest one, the one that uh you can compare with uh X um size of uh popular uh YOLO models. Now here is like a visualization. interesting part about this model I think is uh like anybody who used key points in the past they are very familiar with u like a p if you if you do key point detection uh those models like like yolo they will give you coordinates of your skeleton so for every so skeleton contains multiple anchors and for every anchor you will get x and y and you will also to get like a float uh that represent the confidence of the model. So it tells you like how confident am I to the given key points is uh present and is there. Um so what we did is we um approached this problem slightly differently. Instead of giving you the key point what we give you as um like a spatial con confidence. So what we what we do is we uh give you xy coordinate plus this uh spatial confidence that is uh visualized um as ellipse. Um and the the broader the ellipse is the least confidence uh the model has in location of that point. Um and of course ellipse can be uh you know can he can can have u axises that are uh you know uh dimensionally wise very close to each other or not and that can also represent how confident the model is about the presence of the given point in a specific direction which is also pretty cool. um you can use that information and you can probably plug it into like a common filter downstream to extract even more information about that. Um but most importantly it gives you a lot more uh like a lot more information because before all you get is a float and okay if I do the thresholding at like 0.3 I all I know is that oh here are the points that the model is confident about and okay I don’t know anything more I all I know is like it it might be here but the model doesn’t really know. So here what we have is like um it’s actually calibrated. So I can tell you um like for example here those dotted ones are at uh uh sigma uh two uh and um if if we look at those those that that have like more than sigma 2 those are the the ones that are not visible. Um and also uh if you if you look at the cocoa and you would benchmark that that uh uh and calculate you would learn that if I’m not mistaken uh 41% of the points are always within uh sigma 1 and like 83% of the 86% of the points sorry are uh within sigma 2. Yeah. So that gives you like a a pretty good information about like how confident the model is that the point is actually within um that location. Um and that information is super useful. I mean anybody that’s used some of these key point tracker detection tools knows how often they will sort of you know lose confidence or you’ll drop a point or or or the point will move significantly. And so that’s a that’s a huge improvement here. Yeah. Yeah. So, here here’s an example of I’m I’m just in my my kitchen and I’m just uh kind of like rotating and you can see the mo the moment when I’m kind of like my side is towards the camera is the moment where obviously the model is least confident about the position of the uh of the point and you can see that that okay I mean I I know that like a general location um but I’m not that confident that the location is not that narrow. Um also like interesting if you would go back a few slides um yeah that might be interesting. So here here we there are also those ellipses but we don’t see them because they are so small. So what was also happening that depending on um so maybe maybe let’s say uh differently when you annotate images of people you can be a lot more certain about the location of an eye because it’s very easy to like it’s very easy to see so when you annotate you just click the same with the nose yeah because it’s pretty obvious where it is but for example with heaps or shoulders especially with heaps I would say uh those ellipses are usually a lot uh you know elongated a lot a lot larger. uh and that’s because just way harder to annotate so that there is a a lot higher level of like variance within the training data and you know ultimately that get transferred into your model and when you kind of like train you can see that the mold learns okay like in ter like when when it’s eye very easy for me to locate that I’m pretty sure it’s in this like very very narrow part of an image and with heap is like you know it’s hard to tell. There is like a you know uh different uh posture, different uh different clothing, you know, it’s probably somewhere over there and it learns that from the data. Another very interesting thing is that like I said all of that is learned um uh and that is important because then when you apply this model and you fine-tune it on other data set um we will also learn the distribution from that data set. So it doesn’t matter if if you have uh pose estimation or a very popular use case that I’m uh using key points model to is like to locate characteristic points on like a football field or basketball court. Um it will also learn uh that distribution. Um, an interesting thing that people maybe not know is a lot of those uh keyoint models those popular like open source let’s call them this way uh keyoint model they they almost like hardcode the information about the amount of key points and the distribution of those key points and the level of certainty into their their architecture and their loss function. That happens certainly with uh some popular YOLO models. they they literally add this information like 17 points and eyes are you know usually this uh certain and heaps are usually so all of that is hardcoded into the architecture. So you can imagine that when you apply that uh model later on and you fine-tune it on a different data set that is completely different, all of that information is definitely not helping probably hurting your training and making it harder to train and because our DTR just learns everything from the distribution of data. U then it’s actually useful and you can later on use it during the inference but it also helps during the training. Um and another uh interesting kind of like a property of uh uh our like transformer model it it was working very similarly with detection and segmentation is that when uh you can actually use a different resolution uh as an input for the model. So by default um the extra-large model that we released is uh 576 pixels u square but you can decide that you would like to either lower the input resolution or uh make it larger and once you do it uh this kind of like curve is uh is being generated because you using a different resolution will impact the speed of your model. So that whole curve is created from a single checkpoint um which is located over here for the default configuration. But you can change that configuration during the inference and that can uh either increase your accuracy or lower your accuracy uh and also impact your your training uh your um sorry inference speed. So yeah like with everything it’s it’s all about the data you know uh garbage in garbage out. classic adage still holds true. We’ve also got a special guest here. Please say hi to Fluffy. Fluffy. Fluffy is very interested in computer vision. Yeah, my dog just came back from the from the walk. I I hear in the background. So yeah, he’s also somewhere over there. Very interested in computer vision as well. Um and that’s that’s pretty much it. If you would like to use RFDTR, you can you can use it through our open source package. The QR code that you see right now on the screen will lead to the to the repo. We highly appreciate and stars if you are there. Uh and uh if you want to get even better performance um then yeah I highly encourage you to um to use our Roboflow version of RFDTR because the important thing is um uh our RFDTR comes with NAS neural architecture search but the open source version does not have that. The product version have that. So on average you can usually expect another two maybe 5 map uh boost just by using that um that NAS. I can tell you more about NAS if you want but uh that’s pretty much it. That’s pretty much it. Thank you for all that killer info on RF uh DTOR. It’s uh we saw at CBPR tons of booths were just using this as their demo. There were there were a bunch of them like I recognize that model. I know exactly what you’re doing. Um this was I think before the key point preview but um still really cool product but it was not an open source yet. I uh I was so GPR was kind of like two weekends ago. Uh last weekend I was on another uh conference this time in Europe and tons of tons of people uh told me they’re using our DTR in their product because what is important and I haven’t said that it’s it’s Apache 2. It’s like no strings attached Apache to license which which you know like you said it’s not very very common common in computer vision uh anymore. So it’s pretty yeah these days uh it seems like yeah I would love to hear that too. The it’s a very permissive license. You know Apache 2 is great for uh putting out software that you want people to use. You know you can say go use this for a commercial product. Use this for an open source product. Do do what you will with it. All you got to do is tell people what you used. And I love that. Um so another uh ribflow also had another relatively uh another pretty big release right supervisor had some upgrades. Is that is that the case? Uh supervision I guess. Yeah. Yeah. Supervision. Yeah. Sorry. So uh what happened is that we uh yeah key points was exist. So supervision is our library that we use to power uh our product our demos. So every demo that for those of you who who follow me and aware I blew a lot of demos that’s part of my Jeff and all of those demos are powered by supervision. Um and yeah uh key points were in supervision uh before but you know we’re kind of like the afterthought and because we release uh the preview version of the keyoint mode we we added a tons of new features new annotators new utils you know improved the user experience around this particular part of the library and um we will uh I already know we will release another release around key points soon because uh we built a lot right now with key points and all of those improvements are landing into supervision. So yeah that that was that was the release the it it powers it actually powers our DTR uh package as well. So when you use key points you actually use supervision internally and uh in product you use supervision and in other libraries that we have use supervision. So it’s kind of like a workhorse that we have. Okay. I didn’t know that. That’s good information to have. Um, please scan that QR code at the bottom of the screen, folks. Try out Rooflow today. They’ve got a great free plan. You can try out a bunch of their awesome tooling. Rooflow makes some of the absolute best tools in the business, if not the best tools in the business. And uh, they’re a good group of folks, some really nice, their customer support people are always really cool as well. Um, can’t recommend Rooflow enough. Today’s episode is in fact brought to you by Rooflow. Scan that QR code, sign up. Um, and also OpenCV gets a little kickback if you become a paying customer of Rooflow, which we encourage you to do. So, we think, uh, as soon as you see what Robo Flow can do for you, you’ll want to start paying them for the privilege. And so, uh, when you do that, OpenCV gets a little bit of money back. So, support open source computer vision and get some great tooling as part of the deal. Um, so, uh, Peter, there’s, uh, I saw I’ve seen on LinkedIn quite a few good, uh, sort of third party like demos. I know a lot of people are using, uh, RFD, uh, DTR. Uh, do you can you talk a little bit about some of the cooler, um, uh, sort of third party implementations you’ve seen out there? Like what kind of cool stuff are people using this for in the wild? Um, sure. So, so first of all, um because RDTR is like I said a patch tool, uh that uh makes it uh possible for other libraries to also um other DTR. So, uh it’s now available in uh transformers. So, if you would like to use it for transformers package, it’s possible. Uh we are also working on making that super easy to use RFDTR in mobile uh apps. Um so it’s right now available in uh React Native Exeutor library. So they have all the sizes um uh in different tasks even the key points task available there. So if you build mobile app, you can pretty much um almost like drag and drop your model into your uh into your library. So that’s from the open source perspective. It’s getting a lot more uh broad adoption. Um now when it comes to u industry uh yeah tons tons of people uh use it for all sorts of like smart city use cases. So cases where you would like to know uh how many people how many cars uh are in a given location or moved from point A to point B. Um tons of use cases there. Um sports from what I heard it’s another one. Um a lot of use cases working with CCTV cameras. a lot of use cases working with uh um manufacturing. So you know camera over conveyor and um you need to count objects, you need to make sure that they are intact. um plenty of use cases around like transport uh so uh uh shipping yards um stuff like that that that’s also very popular. So generally any place where uh you would like to build a product and you would like to have a model with open source license um but also um get high speed and high accuracy because that’s important. Um and yeah that’s that’s I that that is pretty much it I think. Yeah. Yeah. Right on. And as I said some cool videos over on uh the LinkedIn account. you’re a great LinkedIn follow because you’re always posting cool demos with a lot of this tech. Um, this this I’m not an

    I’m saying something. Yeah. Yeah. Yeah. It’s a it’s a a careful line you got to walk between uh sometimes you get it’s it’s possible to be too excited about your own work, you know, and like put it out too much, but I think you you do a good job. Um we got a bunch of questions here from the audience. This was uh yeah this was uh generated a lot of yeah generated a lot of discussion here. Um here’s uh let’s see first one here we’ll do the new the most recent one cat 5D dev. Um first they say heart rooflow. We heart roofflow as well. Um but they also ask you mentioned the currently released pose estimation key points would be the biggest available. Are there any plans to add hand or digit key points in the future? So, absolutely. Yes, absolutely. Um, the idea that we have is that we would like to add more sizes first of all. Um, and like I said, like the preview is for us to collect ideas from the community. Uh, look for any improvements that we can do. We already have several um things that we will change for the next release most likely. Uh so you can expect more sizes but you can also expect that we will release uh key points models for different uh use cases. Uh and we’ll have pose estimation and we’ll have hand uh gesture key points as well. um a little bit uh I’m a little bit unsure about the license of that model because those um data sets uh that exist out there for hand gesture have pre u yeah they are okay for research uh but not for enterprise so I’m unsure if we will be able to release those checkpoints under Apache tool like officially it’s not because we don’t want to it’s because we we have hard time to locate locate the data as you said data is the king so it’s pretty pretty important that the license is okay there but we will certainly release um uh checkpoints for uh for hand uh key points and if if we’ll find the data set then that that will be okay uh then that will be a patch to release as well yeah so you can expect all of that to happen the next like weeks months Um, for sure we will release the um the post estimation key points first. All of them. Got it. Sounds like there’s a lot coming down the pike. Uh, looking forward to that. We’ve got a bunch more questions here. Thanks for the question. Cap 5D. Got Eric Feno asks um a similar question. Do you think about doing foot key points? Heel, big toe, small toe. Is that something that’s in the been discussed as well? Uh I’m not sure if so there are two things separate. So there’s like uh what I will be doing and there is what the research team will be doing. Um so my intention is totally to uh to try that. Um there is a cocoa data set that contains this information. So for the reference um the key model that released is 17 key points uh on the body. That means that for example when it comes to hands you only get a key point in your wrist and when it comes to uh your I mean wrist you know uh um shoulder uh and Jesus I am I forgot how how you call this but elbow. Elbow. Thank you very much. I’m not the native speaker and you’re doing great. my hour 11 of working but um elbow and obviously hips and obviously knees and obviously ankles um so all together with face key points that’s 17 but uh there are other data sets out there and very often those data sets I you know have like 21 points per per your hand and another like six points per your foot that allows you to dramatically increase the accuracy of your model especially for like a think I’m I’m thinking always about like sports use cases when I’m doing like football analysis or soccer analysis depending on where you live. Uh it’s pretty hard to do any accurate analysis with like 17 key points because all you know is is where is your ankle but if you have those another six points per foot then you can do pretty pretty accurate analysis. So, uh, coming back to your question, um, I’m not sure if we will do like a official release of that checkpoint, but I personally absolutely intend to fine-tune uh, RDTR on uh, that data set and releasing that checkpoint. Um, yeah, for everybody to just have fun and be able to do cooler stuff. Uh maybe we’ll do like some sort of like a fine-tuning tutorial around this idea. I’m not sure yet, but um I’m certainly super interested in in in doing that. Thank you for that. And we appreciate you uh extending your workday to educate us here. I was just I was just explaining why I cannot speak English anymore. No, it’s okay. Not that I’m complaining or anything. You have a much better excuse than I do. I mean, I’m I’m a native English speaker and I just got up, so Yeah. Yeah. Yeah. You surprised me that you’re in Mexico, but Yeah. Yeah. I love Mexico. Uh, you Mexico. So, we’ve got a few more questions. I’m going to break up the questions here for one moment and do our trivia giveaway. So, folks that are longtime watchers know there’s something we do on every single episode of this show and it’s give away something to you out there in the audience. Today we’re giving away a free OpenCV University course. You can see what courses are on offer by going to opencv.org/university or by scanning the QR code at the bottom of the screen appearing directly underneath my face. That’s opencvuniversity.org opencv.org/university or just scan the QR code. Today’s trivia winner will win the OpenCV University course of their choosing. See all of them that are available there on the website. If you have won in the last couple of months, don’t answer. But do feel free to answer wherever you’re watching. We’re talking Twitch, LinkedIn, Zoom, and YouTube. The first person to answer will win today’s giveaway. Get ready to answer now. So, today we talked about some benchmarks and how uh the new Rooflow release is uh better than uh many of the top stuff out there in what it’s trying to do. One of those was MS Coco. What does MS Cocco stand for? What is the acronym MS CO stand for? Put the whole thing in the chat now and you will win an OpenCV university course.

    I’m watching the chat. Don’t disappoint me today, folks.

    All right, we got one. We got uh Zack Crane, our Ian um for the for the episode. Uh it stands for Microsoft Common Objects in Context. Zack was the first one to get there. Congratulations, Zack Crane watching on Zoom from Iowa. Please send one email to me. That’s Phil opencv.org with the name of the course you would like. Once again, see those courses by scan the QR code at the bottom of the screen and choose one. send me an email about it and we’ll make sure that you get that unlocked in your OpenCV University dashboard. Congrats, Zack. I think this is the first time you’ve won, Zack. You’ve been watching for a while. Uh, love to see it. Thanks for joining us again. And we’ve got a bunch more questions here. So, let’s pop open the question tube once more.

    You guys got that one pretty quick. Uh, we also also Lowi time on YouTube. Cat 5Dev on YouTube. We were you were this close, folks, but you just didn’t quite get there. Um, wow. Yeah, so many questions. Okay. Uh, let’s see. We’ve got Muhammad uh Elbaz on LinkedIn asking, “How many classes can RFDR detect in model object detection? How many classes?” Speaking of cocoa, uh, it’s pre-trained on Cocoa. So it can detect the standard 80 classes that uh you are probably familiar from other open source object detectors. Like I said it’s on it’s it’s kind of like it’s pretty like the checkpoint that you get is trained on cocoa but the knowledge about other things is there in the model. So uh you can still fine- tune it but out of the box 80 class. Okay thanks for that. And we’ve got uh so you covered this question a little bit in the presentation that the talk earlier um but what are the big differences between RFDTR and YOLO? Uh that’s sounds like I’ve not narrowed down to key points only. So I I will I will answer as uh kind of like generally uh not necessarily uh only for key points. So what are the differences? Uh pretty significant ones uh to be honest with you. Uh it starts with the architecture. Uh so y architecture is the convolutional neural network. Um pretty much it’s you know staple of computer vision for the past I don’t even know probably like 10 years or so uh maybe even more than that. um in some way or another. Um and uh RDTR is global flow detection transformer. Uh that means that we we borrow a lot from uh from from what is happening in transformer space right now. Um so that’s the difference like the architecture is the difference. Um that has consequences. Yeah. So uh for example the consequence uh of of that architecture is that connets uh like yellow requires non-max suppression which is this kind of like mechanism that you need to apply at the end to figure out which boxes are duplicates. You know which boxes um kind of like you can kind of like remove and only keep those that are the key ones the most important ones. the transformers architectures don’t need that kind of like they learn which boxes are important and which not. So it’s kind of like a very unified architecture. You can um uh everything happens in a single uh swoop. That means that for example when you export the model uh those models that require additional NMS yeah you exported that model let’s say to onx but you still need to apply normal expression at the end uh when it comes to transformers if you convert to onx that kind of like that mechanism of removing duplicates very much embedded embedded into the model uh of course kind of like on a on a different side of the spectrum Um uh we we made this transformer that is to be extremely fast but it’s still transformer. So uh the model size is uh larger uh compared to comnets. Uh but we are working on like making that not a problem. Um so for example like even though that we have more parameters not all of the parameters are used during the inference um and so on. Uh so so there’s that. Um what is uh maybe one more difference that I already covered but I think it’s very important. Um is how we train. So we use pre-trained backbones like uh Dino uh as a kind of like you know starting point for the model whereas connets just learn from scratch learn from data sets like imageet or koko that means that they only seem those images that are in imageet or koko if you ever look inside koko data set you would be very surprised about the quality of data you know it’s a data set that was released pretty far away in time it was largely crowdsourced. So it is not like enterprise quality and um there are only 80 classes. So the model only learned what was there versus uh we use Dino V2 a model that was trained on massive masses massive data sets very broad data sets. So the model knows a lot about uh the whole world and that uh ends up be very useful when you fine-tune the model because we we just access information that is already in the in the model uh during training. So as a result, for example, when you fine-tune uh RDTR and custom data sets that are like aerial data sets or marine data sets or uh medical data sets, um your RDTR model will score usually several map points higher than Koko. Um yeah, there are probably a lot more differences, but I think those are the most important ones.

    Yeah, thanks for that. Uh, great thorough answer. Um, got a couple more questions and a little bit more time to answer them. Uh, Stefan should stay longer. Usually when there are questions that means that it’s a good sign. So, you know. Yeah. Um, uh, Stefan Schneider on LinkedIn asks, “Since RF DTOR is transformer based, would it benefit from things like visual primitives or is that only a thing for the training stage?” Uh, oh, I think I think that that is the the question that when I’m I’m tapping out um I’m happy to because I know that Stefan uh follows me on on uh LinkedIn. I’m happy to ask this question to our research team and uh give him back the answer. It’s probably above my pay grade. Sorry guys. Understood. That means it’s a good question. Thanks for that one. It’s a good question probably. Yes, but uh not for the open source is probably understood. Understood. Uh we’ve got one from Eric Feno who asks this uncertainty around key points. Is it just standard deviation of the heat map or something else? Yeah. So importantly, there were a key points model in the past that used this kind of like you know uh like a like a heat map to describe uh how certain the those model were uh where the point was like probably notable um one was uh Vos. Um but if you would visualize those heat maps that they uh generated um every point has the same heat map. Uh at least that was the case with uh beat post. So it was not very informative. The only like information you would get is that you know the further away in any direction from that point you would get the least certain about presence of the actual anchor. the model was in our case it’s all learned uh it’s it’s a uh it varies with direction so we also provide you information about how how certain we are about the point being in direction X versus direction Y um it differs between the points so like I said our like a spatial confidence is different for example for eyes than for hips um that wasn’t the case with VO. VO was equally certain or uncertain about uh locations of points um regardless of of the distribution of those points in the data set for example. Um so yeah there there are differences there are differences like I said and we actually don’t predict heat maps we we actually predict like additional parameters and we just then represent uh um them visually as ellipses but um we we don’t like generate heat maps internally. Got it. Thank you for that and thanks for the question Eric. We appreciate it. Uh we’ve got another question here from YouTube. You mentioned that this was not trained with augmentations. Is that possible with Rooflow? Is it possible to apply adversarial training in Rooflow? So uh absolutely I mean you can so when you train RFDTR both in the open source repo and in the pro like product version of the model you can apply augmentations. Uh so question is in in some way important because that wasn’t the case initially in the open source repo. Initially in open source repo we we were like uh only releasing with like the barebone model with no augmentation pipeline. Right now we allow you to perform all popular augmentations available and no augmentations I believe. though like one of those like open source very popular libraries for augmentations. Uh so you can add it. You can add it and it will probably give you some boost. Uh but I can tell you like if you would if you would take Yola model and take our augmentations it will probably drop by 20 map points or something like this. Uh if you add augmentations on top of um RFDTR, you might expect probably like a a percent of map uh boost. So it’s not very crazy. Um but it will of course increase your training time. So if you if you care like primarily about getting accuracy boost and uh less about how long it will train or how much GPU it will burn then you can add it both in open source and product. All right, thank you for that. Um good to know. I I feel like I learn something every time we have Rooflow on the show. Um, we’ve got love love for cricket asks, uh, is this still 2D key points or is there a way to convert to 3D? It’s a 2D key. It’s a 2D key point. We we I I spoke even with the with the ML team. Uh, yeah, for now we we are focusing on there’s like not no immediate uh like direction that we have. Got it. Understood. Um there’s one longer one here from uh Simon who asks uh in the preview that you released, can we retrain the model or are the weights frozen? Um also when fine-tuning the model, what data do we need to provide to the model? Do we need to provide sigas together with key points? Um etc. He was thinking about fine-tuning the model uh for the use case you mentioned uh soccer fields. Yeah. So, uh yeah. Uh so I’m I’m just laughing because I’m I’m I’m literally training model uh for that use case right now because I uh I’m thinking maybe about using that for the tutorial. Uh but long story short, yes, you can uh use this model for fine-tuning. So uh the preview model in the product allows you to fine-tune. the preview model and the open source model um repo allows you to fine-tune. Uh what do you need to have? Uh I I’m not sure I’m pretty sure we um support like Cocoa uh data set for key points. I’m not sure if we support YOLO uh data sets for key points. Um but we would probably you probably need to take a look into the repo, but one one of those formats we support. Um and yet you don’t need to provide any sigas. All of that uh is learned from the data set. Uh so all we need is is a pretty standard representation of keyoint data set. Um and that’s it. Uh although uh bear in mind that was the case with a preview version of a segmentation model. uh it might be the case that uh you won’t be a like that your preview model that you fine-tuned and you got from the preview version of key point then you would need to use that preview architecture all the way and when we do the actual release we will actually release like four new completely brand new checkpoints so the preview model won’t be the X or the L from our final release it will be just like a sep separate completely separate model uh and probably might be the case it will be deprecated uh uh right at the release time. So it’s like you know keep that in mind but we still support that segmentation model is still supported in the actual package uh even the preview version uh but I don’t know there might be some features in the ultimate uh model that we released that that are not in preview I don’t know but it’s totally trainable even right now actually I would super encourage everybody to do that because uh any anything that you notice can help us to to make better. Um, so yeah, do it. Let us know how it went. Yeah, please do. And uh if you try something out from uh RoboFlow today from the show, uh tag OpenCV on on wherever you post it on LinkedIn or wherever and we’ll uh we’ll boost your post for sure. We’ve got one more question here which is from uh Ner Nurmmitic. Will this model be a good choice? not just for object detection, segmentation and classification tasks but also for anticipation of events in long- form videos as well. So important thing to mention we don’t support classification task and I don’t think we will uh do uh I also had a conversation with research team and we just think that there’s no point on creating another classification architecture uh models like restnets or or other models like this are already very good at classification uh so we support object detection segmentation key points uh just just a comment on on that part. And uh when it comes to like a long form um sorry like a you know anticipating events or detecting events in long videos, no I I don’t think so. Uh I mean you can build pipelines um around tasks like this but I think that uh that particular task will be better solved with um VLMs. It highly depends on you know what’s your compute capacity if you want to deploy this kind of like action or event detector um on the edge maybe in that case it it makes sense to have like a very lightweight model like RDTR that can do the heavy lifting and a little bit like a thin thin layer of logic on top of that. um maybe that’s the case but I think like long term um that task will be uh handled by VLMs and and we are actually very much interested about that in Roboflow uh spoiler alert we we we are building like a leaderboard of VLMs uh to measure how well they are uh handling task like event detection uh in videos that’s probably coming uh in the very near future. So um exciting. Yeah. Yeah. Thanks for that. We also just got one more question coming in at the wire here from uh YouTube which is generally transformers are used for LLMs and vision language models. Is RF deter vision only? Why is it benchmarked against CNN based YOLO instead of VLMs like Quenv, Gemma, etc.? Interesting question. Okay, very interesting question. So actually um yeah there there were detection transformers before uh RFDTR. Um

    so uh that let’s let’s let’s let’s say it uh uh right away um the qu the the the reason is scale and speed and also the behavior of the model. So, LLMs or VLMs those are outer aggressive models that means that they generate tokens uh and those tokens then can be uh converted into words or coordinates for bounding boxes. Uh but that is not the behavior of RFDR. RFDR is not out not an auto reggressive model. Um so first of all that’s why we don’t put them in the same category. the all of those models use attention as part of the architecture but there is a certain distinction in how the model behave and how the prediction looks like and um RLMs and VLMs are just in completely different category um because they are reaggressive and also highly impacts their speed so uh most of those models I know that there are right now a little bit different VLMs can do it slightly differently. But most of those models predict one token at a time, which means that even if you would like to detect a single bounding box, you actually need to have like a multiple forward passes to to get the coordinates and the confidence. So we detect all of the bounding boxes all at once just like YOLO models do. uh and why we benchmark RFDTR against uh YOLO models is is because of the speed and accuracy ratio. Yeah. So those are real time uh detection models. So those are the models that you would deploy and expect to get tens or hundreds of FPS uh per second from both Yoro models and from our or detection transformers models. Um so I I think that that summarizes it well. It’s like completely different behavior. The only common part is they they use attention. Uh so that’s why we separate uh them between VLMs and LLMs and why we put them in the same bucket as yellow is because the speed and accuracy uh because they are kind of similar clearly uh targeting the same use cases. So that’s why yeah that totally makes sense to me. Um thank you for the thorough explanation. So that’s all of our questions here. I’m going to take a a brief moment to uh do a quick thank you to all of our sponsors on GitHub. One of the ways that you can support OpenCV is by sponsoring us on GitHub. You can go to github.com/sponsors/opencv and become a sponsor for just n bucks a month. We’ve got 17 of y’all on there right now. And I’m going to thank you individually. Now, we’ve got uh Zashon, we’ve got Alberta Beef. Great, great username. The homies at Rooflow, Bears, Dev, OpenCV bronze members as well, Axel T81, DJ Greenwood, Tac Tealoski, Techman, Dan Dagaru, we’ve got uh Chonx Fres Hugh. I did my best on that one. Please tell me how to actually pronounce that if you’re watching. We’ve got Stefan Sarnv, Luxronic AI, Alexander Voronov, Big Vision LLC, IPOP AI, Comet ML, Alexander Ismolof, and Tala Hussein. Thanks to all of you so very much for your support. You can be as cool as those people and get your names read out here on OpenCV Live by becoming an OpenCV sponsor on GitHub. That’s github.com/sponsors/opencv or just go to the GitHub repository of OpenCV. Click the little heart icon and it’ll take you to the same page. Thanks so much everybody. We really really appreciate you. You are each individually the absolute best. And Peter, you’re also the best. Uh do you have any final thoughts for the audience here before we call it a day? Uh yeah, I’m not sure if I’m the best, but I’ll take the compliments and say thank you. Uh yeah, final thoughts. I mean guys, use use RFDTR. Uh let us know if it works for you. Let us know if anything breaks. Open uh issues on GitHub if anything is wrong or you have ideas on how to improve that. And yeah, let’s let’s keep it rolling. Let’s let’s make this uh open source Apache to license detector as good as possible uh together. So yeah, that’s that’s it. Thanks so much, dude. Um, learned a lot today. We hope you learned a lot today out there as well. One more way you can support OpenCV, become a member of the OpenCV YouTube channel. You can find out how to do that by scanning the QR code or just clicking the button if you’re watching us on YouTube. That’s one of the ways you can support this essential nonprofit open-source software. We will see you next week. Um, same bat time, same bat channel, 9:00 am Pacific time with our guest, Dennis Baldwin of Drone Blocks.

    Excuse me there. Our guest, Dennis Baldwin of Drone Blocks. Drone Blocks is making some really cool stuff for drones. We’re excited to be getting back to talking about some more robotics, especially autonomous vehicles. We hope you’ll join us for that as always, right here on Twitch, on YouTube, on LinkedIn, and on Zoom. Until then, take care of yourselves out there, folks. Take care of somebody else if you can. Use OpenCV5 and have a great day wherever you may be. Adios.

  • Qwen 2.5-VL Tutorial – OpenCV Live 182

    Please in the chat wherever you’re watching on Zoom, Twitch, LinkedIn Live, YouTube, or even for some reason on Facebook, let us know where you are joining from. We love to see the global OpenCV audience chiming in every Thursday morning here on OpenCV Live. I am coming to you today from beautiful 21. Um, and uh, Doc, where you coming from? San Diego, California. And it’s getting the usual places. The usual places.

    I am fixing our Zoom feed real quick. Um, it seems like Okay.

    Yes, that was the problem. Uh, sorry Z. Oh, I can’t apologize to them cuz they’re not involved. They’re not in here yet. But all right, Zoom is now live. We should be good everywhere. Um, I know this is super exciting to watch a guy push buttons you can’t see, but uh we do have an exciting episode for you today. We’ve got our very own Dr. Sat will be giving us a tutorial on the Quinn 2.5 VL multimodal language model from Alibaba Cloud. I think a lot of people out there that aren’t uh clued super clued in to the uh computer vision scene uh as it were may not realize that uh Alibaba has such a massive um uh AI infrastructure within the company at this point. Um, I think they know about them at all. Most out there know that the even cheaper version of Amazon basically

    like knockoff electronics components. They know they’re they’re a massive company and they have and the thing Yeah. And the thing is that they are also, you know, the solutions they have chosen the open source path. So the weights are open source. It’s really uh you know they’re doing a very good job and uh one of their pet peeves is that you know our models are so good but somehow people don’t uh they they don’t reference us in papers right everybody outside the papers right everybody knows that the models are really good for some reason they they don’t get mentioned right so it’s one of those things that um sometimes it is just it is just sad that people don’t the attention that they they deserve. So we will we will give Quen some attention.

    We know all about that here at OpenCV. Something we discuss quite frequently. Um Zoom looks like it’s up everybody on Zoom. Sorry about the late start here today. I am uh not in my wizard tower floating high above Selma in San Francisco, California, but in instead in Tana, Mexico um the happiest place on earth as Crusty the Clown famously called it. And it is that for me I must say every time I come here something else impresses me about it. Um especially the food the food in Tana is fantastic. If you ever get the chance um definitely cross the border and have a visit. It’s um uh a really interesting uh place with awesome food and and extremely cool people. Um but yeah, Zoomers, uh let me know. Let us know where you’re joining from because you missed my little early preamble there. Looks like we’ve got some folks coming in from Nigeria. We’ve got uh Baja. Hey, how you Baja? Uh UAE chiming in here. Um so yeah, let us know where you’re joining from. I know we started a little bit late and we’re also having maybe some small video issues, but it looks like we are in fact live everywhere and so we’ll uh uh get started here. I think Doc, if you’re ready to go, um let’s go ahead and do it. Uh do you want to talk a little bit about OpenCVU at the top? So uh let let me officially start it. Right. Hello everybody. Welcome to OpenCV live. In today’s episode, we are going to discuss Quen 2.5VL. This is a large visual language model where we will discuss what are the capabilities of these model and how to do image captioning, object detection using this model. Very powerful model as you will see. But before we uh start, we also have with us Phil Nelson who is the director of content and creative at OpenCV. He produces the show. If anything goes wrong, it is his fault. Hi Phil.

    Good morning out there everybody. Yes, it’s me. It’s me P I P. I am the co-host with the co-host the second of the second. I am also your plus one and only. It’s Mr. Nelson if you’re nasty but you my dear friends can call me Phil. And I’m here to remind you of a few things that we do on every single episode of this year program. The first of which is a giveaway to you out there in the audience. Stay tuned later on in the episode. I will be asking a trivia question based on the presentation by our illustrious host Dr. Satia Malik. And the very first person to answer that question correctly will win the Open CV University course of their choosing. We’ve got a little bit more to say about OpenCV University in a few minutes. But we’re also taking questions from you in the audience. Use the Q&A button wherever you’re watching in the chat. Just type your question or if you’re on Zoom, hit that Q&A button. save those questions for us in about 40 minutes or so. We’ll do Q&A uh afterwards uh after we do our trivia segment. So stay tuned. We like to alert people to pay attention. Uh if you want that course, watch the episode. Remember what the doc says and you too could win. Um but to start us off here, let’s briefly discuss. We’ve got some uh some some big open CVU news here. I’ll bring the screen share up. Yeah. So the sale has started. We started early. That’s why it’s called early bird special. It is this early bird special is going to end uh tonight. So if you’re interested in OpenCV courses, this is the absolute best time to start. So uh the the thing that I recommend is uh you know it’s a collection. If you’re really serious about a career in AI, then uh the best option uh available for computer vision and AI is our computer vision and deep learning master program. So this is a collection of six courses and it basically has all the courses that we offer and there are two courses out of these six courses there are two courses that are super important. It is fundamentals of computer vision and image processing. This is a Python based course and we also have two deep learning courses. You can do u any one of them deep learning with PyTorch or deep learning with TensorFlow and KAS. And you know these are serious courses, right? These are not like uh your normal courses where you go and uh finish it in 7 days, right? These courses will take you uh solid like each course will take you about 3 months to complete. Uh and but but after you done that, you have a very solid foundation of what uh computer vision is and also about deep learning, right? So 3 and 1/2 months for computer vision course, 3 and 1/2 months for the deep learning course, that’s about 7 months and maybe uh 1 month for uh you know uh just just as a a cushion there. So in 8 months you have a very good shot of calling yourself um an AI expert. In fact, I can say without hesitation that this course is better than 99.9% of the courses out there, including top university courses. So, uh so you know this is a very good deal also right now if you look at uh the pricing 40% off that goes away tonight. So, uh if you’re interested, please uh go and check it out. This is for serious people, right? If you just want to dabble with in in computer vision then uh this may not be the right choice for you but if you are serious about getting expertise in computer vision and AI this is the absolute best program for you. uh in addition to this right so people who go and purchase this uh program during this webinar uh and and you can let me know in the chat section you purchase the webinar uh just you know give a screenshot or something uh or or just send us an email at coursesopencv.org board, I will send you a special gift, right? This is a webinar only special and this is basically a book by Frans. This uh of $48 value and uh this is this book is called deep learning with Python and we’ll give you this book for free, right? Uh but you have to purchase this is not included in this um in this webinar uh in this sale Labor Day sale but I have this very special webinar offer for you. If you uh if you purchase it during this webinar before this webinar is over and send send us an email at courses openc.org you will get that book included in this uh in this program. All right so uh let’s get started. Uh, we have a very exciting uh Oh, actually, sweet doc. Yeah. Do you have other things to cover before we get started? No, I think we should just get going here. I’m going to uh put take myself off cam and uh turn the show over to you, Doc. All right. So, uh let me sh uh let me share my screen. How do I share my screen?

    Okay, right here. share screen and we are going to start with.

    Can you guys see my screen?

    Uh, can people see my screen?

    Um, yeah, you’re all good, doc. Sorry about that. I was I was muted out for a sec. I’m having some AV difficulties here, but go for it. All right, no problem. So, okay. So, Quen 2.5VL. First of all, what is a vision language model? Right. Now, um a vision language model is basically uh a a model, right? We we some of you may know or most of you probably know what is a large language model. A large language model is an AI uh model that has been trained on internet scale uh data, right? So we basically took all the all the uh data from the internet uh all the text data from the internet and created a single AI, right? AI model that is so large uh that it it basically encodes all the information all the world information and you’re able to answer um talk to it, right? So uh that’s a large language model but of course we know that uh you know text can go only so far we also need to include images and videos and that’s where vision language model comes in. So you know if you look at traditional computer vision uh algorithms or neural networks like image classification let’s say you want uh you want to build an image classification model you give it an input and the neural network basically would tell you what the output is right it could in a standard classification problem the input is a dog and or input is an image and the output is a class label right and you have this neural network that you train and for training this neural network. You take several examples of dogs, cats, horses, whatever you want, whatever classes you want and uh you basically train the neural network. You have the ground truth also. Uh you also know which image contains uh what, right? So image of a dog, you have the label dog, image of a cat, you have the lab label cat and so on and so forth. And then you train uh a uh you know a neural network uh usually an image class uh class classification network like let’s say ResNet 50 and you it will be able to tell you information about those classes only right so let’s say you had 50 classes in your training set you would be able to get 50 classes you know you’ll be if if there is an image of that class it will be able to classify that okay this is a cat dog etc but this is a very limited way uh because it doesn’t know anything outside that class, right? And just imagine um if you give it let’s say there were no elephants in the class and you gave it u image of an elephant, you it would either classify it as one of those uh other uh 50 classes or you could have a catchall, you know, uh unknown class and that that’s the best it can do, right? it can cannot automatically uh figure out that it’s an elephant. However, if you had trained a visual language model, right, you u and a lot of these models are trained on image caption pairs. So they go on the internet, they find out, okay, this is an image, there’s a caption. So this vision language model because the language part is there, it gets world knowledge from the language part. It is able to infer what is what an elephant is. Even though uh you had never explicitly trained it to recognize elephants, the vision language model would be able to recognize elephants because it has been trained on internet scale uh image caption pairs. Right? So that’s the basic difference between u VLMs and uh traditional uh computer vision techniques like uh CNN’s. Now uh using CNN’s right the normal image classification uh techniques uh both worlds right so vision systems when we uh when we analyze an image using let’s say a convolutional neural network we get features which learn which know a lot about that particular image right so we learn those features and that’s the vision part so you are able to very closely analyze what is inside this image but at the same time the language part because it has been trained on world knowledge from the internet uh you’re combining that world knowledge with local information so the vision part is getting local features and this LLM is basically helping you uh get global context so you are able to do so many things which were completely it’s it almost looks like science fiction now right one uh big difference between traditional methods, right? Uh or it’s it’s funny I’m calling them traditional methods now. Uh but they are pretty new techniques, right? U so let’s say image traditional methods of of so long ago like 5 years. So u now think about image classification when we say image classification uh the input and the output they are constant right and only the neural network architecture changes. When I say that ResNet 50 is an image classification network, you would immediately know what is the input and what is the output. Okay. Similarly, when when I say um okay, we train an object detector, you know that the input is an image and the output is a collection of bounding boxes with labels. So there is no ambiguity in what an object detector does uh or what an image classification network does uh or a segmentation model does. the input and output. As soon as I say that I have an object detected, you are very sure what it actually means, right? But that’s not the case with VLMs. The VLMs, they could mean many different things, right? So, um for example, uh a few weeks back we had done um uh uh we we we had done an episode on uh on clip. So the architecture of clip the inputs that were given to clip the output clip gives is completely different from quen 2.5 VL they are completely except for the word VLM that they are both VLMs they are both visual language models there is no similarity between the two uh right the so this is something that you have to uh bear in mind that this is not like when when people talk about VLMs they can actually talk about many different things right. The only thing constant among them is that the input is um image plus uh you know sometimes even the uh input is not um is not the same. But if you look at modern uh you know the more capable VLMs usually you have an input which is image plus text or video plus text and the output is usually text right it can also be an image but for simplicity we can just keep it text right and so today we are going to talk about uh quen 2.5 VL uh now even though in this diagram I say that okay we can have image and uh text usually Everything is encoded in a prompt, right? So the prompt encodes the image, the text, the video, whatever you want to give it, it is encoded in a prompt and then we have the uh visual language model in between and the output is text. So we’ll show you the exact syntax how it is done um in uh using Jupyter notebooks. But let’s first look at the capability, right? What is this visual language model capable of? Uh what you’re looking at this input, right? This input has four images in the same image and you can ask quen 2.5VL what are these attractions right we are not giving one attraction at a time this is a complex image it has four uh different images uh embedded in it and u we are saying that please give the names in Chinese and English okay and if you do this it will output this kind of thing it would say the great pyramid of Giza and uh you know uh words in Chinese which I’m assuming are the great pyramids of Giza and uh similarly right uh it would say the top right uh we have the great wall of China bottom left the statue of liberty top right uh the terraota army now think about it we did not there was no training involved the model is so capable that you give this complex image it figures out the structure the layout of the image that it’s not a single image but a collection of four images And then it is able to say uh you know what are the various things right what are the various attracts. So this is a very powerful capability which uh you know just a few few years back uh it was it was not possible. In fact a year or two back it was not possible. Um now you may be thinking that oh I’m going to use VLMs for everything now because why would I use uh tradition uh why would I use image classification or object detector when I can just use throw a VLM at it. The problem is these are very heavy models, right? You cannot use them uh you can use some of these models on the edge, right? But they require uh they require compute and things that can be done. So don’t use a very big hammer to solve a small problem, right? Use uh the right tool uh for the right task. So there are things where you just don’t have enough data. If you have no data at all, then start with a VLM. And uh if you have data and you have to you have to solve a very specific task for example if you are uh if you’re building something for a manufacturing u inspection in a manufacturing plant then uh you know the kinds of defects right um and the VLM may not have that level of knowledge right that okay this is a defect because who knows right that manu that manufactured plant only you know that a defect looks like this and you may not have many examples of that defect, right? So, you are much better off training um in an an image classification uh network and solving that problem. So, use the right hammer for the task, right? The right tool for the task. Don’t uh you know uh because these things would be very expensive to run if you uh actually use them. Okay, so here’s another example uh what uh Quen 2.5VL is capable of doing. you’re looking at um a picture okay and you’re asking the VLM detect all motorcyclists in the image and return their locations in the form of the coordinates and then we are giving it the format right we want bounding boxes 2D this is the format it has a label and it has a sublabel the label is motorcyclist the sublabel is wearing helmet uh or not or not wearing helmet okay and if you run it on this particular image you will see this output right. So it has you can see that all the all the uh all the motorcyclists have been identified and um and it also you know the the coordinates have been identified and it says not wearing helmet, not wearing helmet, not wearing helmet and then wearing helmet. You would expect you know uh in in a street scene most of them would be wearing helmets and then there would be one which would be not wearing helmet but um I’m guessing that this is a scene from Asia. In fact when I was growing up they would um nobody nobody wore helmet uh while riding bikes. Um so it’s it’s just funny that you would not see this kind of scene in the in the US. Um but yes uh it it did find all the classes as well as uh the subclasses right and now you can use this model. Uh important thing here is you never trained this model on uh on helmet versus non not helmet. You never trained it explicitly to detect uh bikes. Right? So it has figured out by uh you know by the training process that these kinds of things are helmets, these kinds of things are bikes and it also knows how to localize uh this information. Right? So it is you know fascinating what it is able to do. It is almost like science fiction. One uh one thing I would like to point out is that some of these examples have been run uh the quen 2.5VL comes in many different flavors. So the biggest model uh was used to produce these examples. So sometimes when you use a small model which is which you can do on Google Collab, you may not get the same quality results because you know it’s a small model not as big uh for which these results were generated. U it also understands because of the language part it has um it has a sense of what uh the social structure is right. There are a lot of uh you know when when we uh as humans when we interact with other people there are a lot of you know social information that we absorb from the environment uh which which doesn’t need to be explicitly taught to us right the same thing happens with because of the scale at which these things are operating they understand uh the context right the social context also for example in this case uh you are given an image and it is said you know you you ask it what is which person is as a uh you know acting bravely in this case. There’s a lot going on in this image, right? It needs to know that uh first of all, it needs to completely understand that okay, there is a knife, there is a person, he’s threatening a woman and there is another person who is trying to save, right? But not only that, it also uh needs to know this the idea what bravery is, right? What courage is, it needs to have absorbed that idea as well, right? It is not just that you can locate all these things, right? That is that is a mechanical process, right? that is um that is doable right you can understand that but to identify that in this context this is the person who is acting bravely that is quite something that has it gives you a glimpse of intelligence right not AGI or anything but still a glimpse of intelligence however it makes this decision you may see say that oh it’s just doing some calculation yes our brains are also doing similar calculation but you can see that there is a glimpse of uh intelligence here. Okay. And uh another uh task which it is very good at in fact in our consulting we use uh quen 2.5 VL for u for document when whenever there’s a complex document parsing um and uh we have to use VLMs in that case u 2.5 VLM is our VLM of choice. So it’s very good at document parsing. You can in fact you know in this example you can uh say that okay convert this document into HTML format right that’s something but you can also say give me the summary of figure one right or what does figure 1 say right or where is this figure which does this right so you can have very complex uh document analysis that this thing does and it can also structure your document uh PDF especially it can structure your PDF extract information from the PDF very nicely. So if you’re thinking about building uh a document parser uh this may be uh a very good option. So um so this this is very interesting and then we uh it can also analyze videos. Okay. So this is a video u which uh where there you know they they want to localize particular events. So what’s happening? You can basically say uh put in uh this video and say that localize a series of active activity events in the video. Output this uh start and the time stamp for each um for each right. So let me show you. I think I’ll have to restart my sharing because it is on a different um uh it is on a go for right. We’ll make it work. Yeah. One second.

    Let me share my entire screen instead of just Mhm. entire screen. I will share screen two. Share. Okay. So you can see that now, right? So in this video, right? After you give this prompt, it goes and says that okay, yep. At start time 21 seconds, a person removes a piece of meat from its packaging and cuts off the fat. So, okay, let me go to the video. Where is the video? Yeah, right here. So, let’s go to 21 seconds.

    So, you can see at about 21 seconds, uh, let’s let’s start again. Sorry. So the person is removing 19 seconds the person is removing 21 seconds they removed it put it here and then they start cutting the meat right and the fat uh for my vegetarian friends uh I’m sorry u and for my uh non-vegetarian friends who are salivating I’m sorry as well uh but yeah so uh so basically it is able to identify uh various events in uh in the video, right? And it says that uh at at location 50, the person is seasoning the meat with salt and pepper. And if you go and look at about 50, let’s go here. So, it’s not actually salt and pepper. It’s uh oregano or some other u you can forgive it, right? At least it is a kind of spice that they are putting it on the on the on the meat. So these are the kinds of very sophisticated long video analysis that you can do with Quen. You don’t you don’t do much cooking, do you doc? Why is it salt and pepper?

    Yeah, I’m just just wondering. Just wondering. Well, because I I’m not using the uh the correct language of cooking.

    All right. So, uh video analysis, right? So it it does a very uh excellent video analysis. Now bear in mind that these results uh you will get this quality result if you’re using the biggest model right as the size of the model comes down you should pair down your expectation also. Okay um they uh it can also do sorry okay and it can also do agent action. Now in the context of VLMs uh agent usually means when you want to do something uh when you want to control your laptop screen or your mobile screen using uh using the outputs right you basically give it you take a screenshot of your laptop right um or your or your mobile uh screenshot give it to the VLM it will automatically figure out what the layout is and what you know what are the various ious apps you have and what are the various things going on and let’s say you start a browser it will know that oh you have you are in a browser just by looking at that image and it will be able to do the next task so you can say that oh make a reservation for me um at my favorite restaurant right whatever the restaurant is u so it will go uh to the browser it’s going to uh start the browser and then look for the favorite restaurant go near your location etc etc so that kind of thing is called agent action and uh this is able to do that as well. So uh okay so that’s let’s look at the usage right how do you actually use the VLM the quen 2.5 VL it is u you interact with it using a chat interface right the good thing is that you can chat with it several times you can give it the context um there are uh the chat looks like this right hopefully most people are familiar with uh JSON format right JSON is you know uh text format where uh you can you can pass structured data uh using this format. So here we first uh you know the chat interface looks something like this. Uh it is called chat.comp completion and we’ll see uh very specific examples. Uh you have the model you say okay I want this particular model and this model comes in many different flavors. This is a 7 billion model. You can also have the three billion uh parameter model and so on and so forth and you know based on how many parameters it has the quality will be different. Then we this is the crux right the main thing is the message that you’re going to pass to uh to Quen. There are two kinds uh of messages actually there are three kinds of messages but uh let’s let’s for for this we can just focus on uh one or two. So the first one is called the system right the system prompt is uh telling it at a high level what it is right it’s giving it an identity that you are an a helpful assistant right um and this is the kind of me the kinds of things you put in the system prompt is something that you would ask again and again. Okay. So, uh for example, when we do object detection, the system prompt could be you are uh you are an object detector that returns the response in a JSON format and then you can uh you can you can say what the format is the bounding box format etc. So once you set these things in the system prompt you don’t have to go and change uh say that again and again in the user prompt right. So user prompt is usually reserved for things that you want to get done immediately. Right? So the system gives it an identity and the user prompt is telling you know it is requesting that I want uh something done. So here you can see that the user prompt we are passing in an image URL and there are different types also the content that uh gets passed uh it can be an array of different things and you can you can this array could be many different sizes right you can pass in one image two images multiple images in any uh any sequence you want okay uh but uh when you’re passing image you can pass it an image URL and then you can also pass in a text. Okay. U so here you can say that okay this is the URL and all I want to know is what’s inside the image. So what is the text in this image? So you can send like that. You can see how simple the user interface is right using just this prompt right uh you can you can uh make it work. No other knowledge is required. So system as I had mentioned that it sets global behavior right and uh you you say things like you are an assistant that answers briefly returns currency in USD right let’s say you are doing receipt analysis right you got a receipt and uh you want to uh even if it is uh let’s say um a Japanese receipt you want the output to be in USD format right uh dollar format you can tell it in the beginning and it knows right then the system prompt sets the stage it gives it an identity and then you can say that users and uses JSON where a user asks for structured data. So um here’s another example of user right I already mentioned what it does so we’ll we’ll just uh skip over it but in case of the receipt example you can pass in the receipt and say how much did the latte cost give the answer as JSON this is not required because we have already put it in uh the uh system prompt uh but it doesn’t hurt if you’re asking the same thing uh again and it will give you the answer even if the receipt is Japanese. It is going to give you the answer in USD. Right? So very simple, right? You don’t you didn’t train the model. You didn’t do anything. This you took this model and uh just passed a prompt and it is give you able to give you an answer. And you can do this right today with uh you know uh on on your laptop uh as long as it has u enough uh GPU. I mean it will still work but it may be slow if you do not have a GPU but you can definitely do it in Google Collab and I’ll show you uh how it is done. Okay. So model card maybe maybe send the prompt to go go make a sandwich go get a some coffee or something when you come back it’ll be done right. Uh so uh this is you know a 3.5 3 billion parameter the uh model that we are going to show in our demo. It is a three billion parameter model and it is multimodal. Whenever people talk about multimodal in uh in in AI models it just means that it is capable of taking more than text right so text plus image or video that’s multimodal as well right and if it sometimes uh you know more sophisticated models would take in speech as well. So that’s uh truly multimodal uh produced by Alibaba uh cloud team which is also called the Quen team. The quen uh license uh you have to be a little careful about it is uh you know it’s a research license for non-commercial uh purposes u and the 3 billion and the 72 billion uh sizes they are not Apache 2 right but if you want to research it you want to um you can use it for a lot of things right uh when when you’re building a model to understand the model so that’s that’s good all Right. Uh and the weights are open source, right? So which means that you can actually use it in your uh notebooks etc. Now let’s start with image captioning. You you guys are able to see this um screen, right?

    Um it should be showing a notebook now. See the we see the uh Google collab notebook. All right. Perfect. So let’s uh start with one example, right? We will start with um let me make sure that I I have finished. Yeah. Oh no, I had not finished the Okay, sorry. Uh let’s let’s keep going. Um one of the things that people may be thinking whether whether it will fit in your uh in your GPU or not. So quen 2.5 VL 3 billion right this is going to take about 3.5 to 7 billion um you know uh sorry 3.5 to 7GB which means that an RTX 3050 should be enough and you know the slides will show you um all the different categories that it will you know if you’re using a bigger model you basically need a GPU with a bigger memory. uh this these slides and everything right all this information is in our uh free VLM boot camp as well right so these slides are there you can register for our free VLM boot camp just search for opencv VLM boot camp and you will uh you will find uh this all this uh you know you can register for the free course and you should be able to uh start it or you can scan the QR code that’s up on your screen right now if you’re watching this show live and find it on the OpenCV universe city homepage navigation.

    Yeah. Yeah. In the interest of time, actually I’m uh let me just uh skip some of this very quickly. It can do image captioning, object detection, we’ll see this as well. Universal recognition, which we saw an example of again, object detection, different type. So, uh we have pretty much covered this. Okay. Um okay. So, let’s uh let’s go look at image captioning. Here we are going to use a Google Collab and I’m going to use a 2.5 uh you know uh so if you’re using Google Collab for best results right you can go here and check what is the runtime you’re using so the runtime if you’re using uh collab pro plus or collab pro you will notice that uh you have a GPU option right A100 is the most powerful GPU GPU here uh followed Why I’m guessing I uh I always get confused between the two. I think u the T4 is the second powerful and L4 is the uh the least one but you need to use you should use a GPU if you have access to it otherwise things will be slower right so you can uh go and use this runtime. Uh another thing about collabs or any notebook in general is usually you know many people don’t use this feature. There is uh this table of content if the notebook is prepared properly which in our case uh we take great care to prepare the notebook. Uh you will see this table of content uh on the side which makes it very easy to go and uh navigate the notebook. All right. So the very first thing we need to do is install um Quenv utilities. Right. So we pip install this and uh we should be uh good. Uh one other thing is that 2.5 VL it’s going to fit in uh it it consumes about 6 GB of RAM in float 16 format and uh I’ll show you how uh to use it in float 16 format as well. Uh so you know most GPUs 3090 and above should be able to handle this. We start with uh standard imports. uh not going over these. Most of the people on this webinar should be a uh should should know most of these right. Uh the only thing of interest which can be new to you is uh we are importing right the transformers. So transformers is this library by hugging face which is an excellent uh package for Python package which has access you know using this you have access to pretty much all the you know large language models regular models all kinds of models uh in this in this package uh and it’s very nicely done so that you don’t have to do repetitive tasks again and again u all the things all the gotcha things where uh you suppose you did not do the right pre-processing etc is taken care for you. It does under the hood uh it does a very good job. So from there we are going to get uh this uh quen uh 2.5 VL conditional generation and uh we are you know this is a package we are importing and also autoprocessor. Now autoprocessor basically does um pre-processing. Let’s say you have text. The text you know large language model is not going to take text as input. So you need to tokenize it and then send it to u uh the large language model. Now tokenization as we had covered it in the previous uh previous episode while discussing clip it is basically converting uh text to a numerical representation. A token is roughly speaking roughly speaking it’s a word. Okay. So you can pass uh if you have a string you can convert it to numerical values for pretty much every word you will get one number and that process is called tokenization. Now you don’t have to do it because you have the autoprocessor which will do it for you. All right. And then we have you know standard imaging and visualization. And then we uh this one is the important one. uh we want to download quenv utils and we get the process vision info. So this basically does the post-processing uh uh this basically does the post-processing of the model. The model output may not be uh in the format that you want because um you know so so we can do some post-processing on this to clean up. All right. So the very first thing we need to do is uh we need to set the device. If if a GPU is available then we use the GPU and CUDA basically is u if you it’s specifically Nvidia GPU. CUDA is the library that uh is run on Nvidia GPUs and when whenever we want to refer to Nvidia GPUs we say CUDA right and if CUDA is not available if an Nvidia GPU is not available then use the CPU. Okay. Um and then we uh basically tell the model id this is you know 2.5 e 3 billion uh instruct uh for people who do not know whenever there is uh a model has the name instruct in it means that it is uh it is a chat-like interface right you can interact with the model in uh in a chat format. If the instruct is not there in uh a large language model, it means that it is a raw model, right? It is u most of the time you would not download those models. Those are for researchers who want to use the raw output and not uh this additional instruct model. Okay. Then we set up the model. We extract the model using this uh quen 2.5 for conditional generation dot from pre-trained. So look at this. how easy things are these days, right? You basically can extract this model uh by all the tools available. So you give it the model ID. Not only that, you also say that um I want to automatically figure out um what kind of data type this model should be in. Usually when you get the model that would be in float 32 format, right? But you don’t want uh things in float 32 uh because there are float 16 FP16 which will run much faster on the GPU without much loss in uh accuracy right and when you see auto it knows okay uh FP16 works very well on GPUs but pretty much you know you won’t see much uh much in in CPUs right basically it gets converted if you even if you have an FP16 uh data type for CPUs it gets converted ed to FP32. Uh so so this one does a good job of already downloading the model u in in the right format right in the FP16 model format if you have a GPU uh it also there is something called device map when you have multiple devices or the model is very large it uh first tries to use the GPU and if it cannot fit all the layers in the GPU it can use the CPU etc also so uh you just set it to auto so that it uh does the best job uh with the available hardware then we need the uh uh processor as I said that you need to do some pre-processing and u again you know the the syntax is so easy sometimes I feel that everything has been made so easy for um for the next generation uh how will they know what it takes to write write difficult code uh but anyways so you get the processor uh out and We will use this processor to prep-process the data before sending it to the model. And finally uh when you do model device, it does everything needed. It downloads the model as well as sends it to the device. Right? So this model needs to reside on the GPU. Um and and it does all of that thing. Um and you can see here uh it is downloading. Oops. So it’s uh it’s downloading all these models. uh not this model, sorry, all these files which are part of the same model. And here you can see that um there are you know uh there are two files model 0 uh 01 uhsafe tensors and 02.safe tensors and it is going to download both of them and use it in the load it in the right format. You don’t have to worry about uh any of this uh how how exactly it is done. Okay. Next we are ready to upload an image to this model and do some inference right we are uh we want to do image captioning. So we get an image uh this is we are downloading it from the URL and here we are opening the image and for for people uh this you know if you have programmed in Python this may must be familiar code we basically request basically downloads it’s used to download the URL but this downloaded uh you know data is not in the right format it it doesn’t look like a file system to the uh to the image uh you know uh opener So this io.bbyte uh bytes io basically it makes the data look like a file system as if you’re uploading from the if you’re loading the image from the file system and you can o open the image after you open the image. Shout out to the uh Python request library. One of the most useful libraries I think in in all of Pythonom. Yeah. Yeah. Yeah. That’s that’s absolutely true. Uh yeah. So you can download it in one line of code. Oh, one one other thing that we are doing right here is converting it to RGB format automatically so that um so that you you’re not um you know uh you don’t have you don’t get the image in BGR format and then you have to do it again right um okay so this is the image we have uh downloaded and we want to caption this image now as I me mentioned we have to create a message right and here we don’t even need the system prompt Right? We can just use it’s a simple problem. Uh we don’t even need to provide any system prompt and everything can be done using the user prompt. So message we assign a role to it which is user and the content we want to pass in the image and we want to pass in the text describe this image. That is it right? You uh that’s that’s what a caption is. It’s a description of the image. So we are instructing it to uh describe the image. Now let’s see uh you know um h oh yeah oh so this is sorry I was trying to think why did I write this code twice. This is part of the uh markdown and this is the actual code. Uh so you run this you store this in messages. And now uh now let’s let’s look at what the chat template looks like right u the a chat template is basically we have written everything in JSON format but this format is not something u the model recognizes u it it has a particular chat uh template right and we have to use this uh chat template and the template has things like right when it sees uh in the template it uh in the string it sees image start ro content etc. So you need to send it in this format. Um fortunately we don’t need to uh you know construct this manually. But just to give you an idea uh the JSON format that we created is not the one that would be used uh by the model. So we need to we need to change the format. Uh we still want to input in JSON format because that’s convenient but we want some code that will automatically change it. Right? So you can see here uh we do processor and processor has everything done for you right you don’t have to um you don’t have to worry about it and the the fun part with this proc processor is that when you you basically get it automatically right based on the on the model that you have specified. Uh so when you’re using transformer package you can pretty much use the processor in similar ways without knowing uh exactly how it is doing things right. So it’s very consistently done. So you apply the chat chat template we you pass in the uh message you say tok to tokenize is false and the reason we don’t want to tokenize it right now is in the next I I’ll show you in the next uh you know when we actually send it uh we we’ll turn on tokenization but for now we don’t want to show turn on tokenization I want to see if we if you turn on to tokenization we won’t be able to print these things would all look like numbers okay um add generation prompt right what is the generation prompt that also uh will be added and we are going to print it right. So you can see that image start system you are a helpful assistant. We did not explicitly say it but there is a system prompt which uh it automatically appends it. It knows that okay you did not put in the system prompt. I’m going to put in the system prompt myself. Okay. And then we have uh this next one which says uh you know describe the image. So we are uh sending all this information. So this you know you don’t need to know exactly what is the chat template. Um it may be useful for debugging if you have special characters or something can go wrong. Uh but most of the time you’re working with this kind of a JSON uh message. But it’s good to know that this is the input uh that uh that the model is expecting. Okay. Uh so now we have uh you know the uh basically we are using process vision info. This is a utility that basically if you look at the comment it says it walks through every message. It finds all the instances of image and video and then it applies uh quen visual pre-processing to ensure that the image is in the right format that pilimage uh uh format and then it returns two parallel lists right and uh of image inputs video inputs and then uh also you know u it pre-processes both the text as well as the image oh sorry this one is it pre-processes only the vision part Right. So uh you get image inputs and video inputs. It extracts from the message that you had created. And finally we are ready. Right. Uh we are going to pass in this is the input that is going to go to the model. Uh so we use the uh processor. We get the text prompt. Right? So this one is uh basically the reason for these brackets is that it’s a one element right? It is only one text prompt that we are sending. And then we have the image inputs. We send in the video inputs. Uh padding don’t worry about it. You know, set the padding to true. Right? It basically u yeah let me let me skip this thing. Uh set the padding to true. It is some uh little detail that uh it’s not very important here. And then we have uh return tensor in pt. So transformer it’s a python package and it will work with uh tensorflow as well uh as well. So we have to explicitly tell it to use uh pytorch right. So the tensors that it returns should be in pytorch format. Uh for people who are familiar with tens uh you know uh tensorflow you would know and and pytorch both you would know that the two formats are slightly different right. uh the in in in PyTorch the bat size comes first. The first element is the bat size followed by the number of colors followed by uh the width and height. Uh in TensorFlow the format is slightly different. Uh I cannot remember exactly. I have not used TensorFlow for a while but uh I think the uh it is height width and then channel and then finally the bat size. I cannot I I may be wrong on that. All right. So uh now we want to use okay once you have done this right this input you’re passing it to the model right uh to the device sorry and uh this device we have already set to GPU etc. So we had already passed the we had already transferred our model uh to the device but here we are going to pass the input we are going to uh push the input. Um next we have torch.nograd No grad whenever you’re using in PyTorch whenever we are using some sort of um whenever we want to use inference right we the model we want to make sure that the model is not um not generating all the gradients. So to make things faster this narrat is used uh otherwise it’s going to calculate the gradients and gradient calculation is necessary when you’re doing training but during inference when you’re just using the model that calculation is not necessary and it can unnecessarily slow down the model and uh so basically you uh you wrap it in torch.nograd and say model.generate generate and you pass in the inputs and you also say max new tokens is 64. So this gives you a sense of uh you know uh this caps the output to 64 tokens. Okay. Um so the caption would be you can say roughly 64 words maximum. Now just to give you an idea what these tokens look like, we can do inputs. Uh if you look at these input id zero, this is the token, right? Uh these numbers are basically tokens. Okay? And you will see um so they are they are fixed length tokens which means that after if they have used all the tokens they will start padding it, right? The padding equals true is basically it says that pad uh with with the with the same value which represents the end of the string uh if so that the length is the same. Okay. Um so okay so this this is just the input and now uh we are going to do batch decode right we uh the the output that we get right the generated ids that we get it is not in a format that we can read. So we want to convert this um we want to convert this to a caption that we can actually read. Okay. And uh you know there could be some special tokens that we want to skip. And if you look at this uh after all this thing is done uh and we are printing this this is the output you receive. The image depicts a serene and picture picturesque scene of a white dog sitting on a stone pathway near a stunning lake. The lake has crystal clear you know uh sounds very reasonable. So you can see uh very few lines of code and you are able to do something as sophisticated as image captioning. I was planning to go over the object detection notebook as well but I think we are uh we are running um running short short on time so we will skip that. But if you register for this course, you will be able to uh see the object detection notebook also. We cover that in the free course. So uh just do a Google search on u on opencv.org. Uh Google search on opencvl boot camp and you should be able to get it. Maybe if people really want it, we’ll do a whole we’ll do a whole episode on that. Uh yeah, that’s another option. are really interested, send it send an email to [email protected] and request it. Yeah. Um, all right. So, let me stop sharing my screen.

    If anybody uh purchased the course in the last uh you know during this hour, please send us an email at coursesopencv.org. org so that we can send you uh the additional free gift which is a book by Francois uh deep learning by Python. It is not included in the current um thing you have. Uh I also want to say you know we also have a webinar um that we will do on uh let me share my screen. So we are running this um on the last day of our um on the last day of our um of this you know uh this event uh this sale event we will have uh this webinar that starts uh at 5:00 p.m. specific time and we will go over you know what the courses are everything we will go over this and you will also learn you know why why should you do uh computer vision now all sorts of things so sign up for this webinar at opencv.org/weinar org/webinar uh and uh you know uh this is the last day of the sale event. So we will basically be there for as long as it is necessary for to answer all the questions. Right. So uh yeah 5:00 p.m. Pacific time September 2nd 2025 uh sign up for this webinar. I’ll be there for uh I plan to be there till the end of uh the sale event which is midnight Pacific time. So we’ll start at 5:00 p p.m. Pacific time. Uh I’ll be there throughout we will be talking about many different things you know how to start a career but also about uh various other things right how do you actually do um how do you actually in in a real world scenario how do you go from uh various steps right model deployment this and that. So many of these things we will cover in those uh in that um in that webinar as well. So opencv.org/webinar.

    Yes. And sign up. Um these things are free folks. It’s it’s amazing that they’re still free. Maybe uh Satia will start charging for these at some point. But that that day has not yet arrived. Um yeah and full disclosure full disclosure uh you know this uh this is the last day of u the sale right so there will be selling involved right there is information involved but there is also selling involved um I’m uh I just want to be upfront that it is not uh this kind of webinar which is only educational there it is educational plus selling uh you have been warned Hey, you know, we we try to be honest with you folks out there. So, uh, and speaking of honesty and speaking of free stuff, I think now is a great time to do our giveaway. So, I’m going to go ahead and in fact bring our chat up on the stream here. Um, so, uh, unfortunately I can’t put our Zoom chat up on the screen, but I am monitoring it in a separate window. Um, okay. So, the way this works is I’m going to ask a trivia question based on today’s presentation and the very first person to answer that question correctly in the chat either on LinkedIn, YouTube, Twitch, um, or on Zoom. Also, Satia, there’s a there’s a typo in your in your link there. Um, says webinar, not webinar.

    Um the winner. Thanks for that chat. I appreciate it. Yeah. Um we will uh so the the way this works is I’m gonna ask a trivia question based on today’s presentation and the very first person to answer that correctly will win the OpenCV university course of their choosing. You can go to opencv.org/university

    to see what courses are on offer. or you can scan the QR code that I’ve just put up on the screen right under our faces here if you’re watching on the live stream. Transform your career with computer vision deep learning and OpenCV courses at OpenCV University. Go ahead and scan that QR code. I’m going to remove the chat here for a moment and uh talk a little bit about our sponsors for today’s episode. But today’s episode is brought to you by OpenCV University as uh Doc talked about a little bit earlier. Um there is in fact a fantastic sale going on on OpenCVU. Um you can go to opencv.org/university

    to get the Labor Day sale for just 17 hours 22 minutes and 52 seconds longer using code labor 40. That’s labor 40 for 40% off all OpenCV programs and courses. That is a huge chunk of savings. Don’t miss out on it. That’s opencv.org/university.

    This episode is al also brought to you by and partners. OpenCV is a nonprofit organization that puts out free open-source software and as such we depend on the support of organizations just like the ones on your screen right now. Big shout outs to ARM, Qualcomm, RunPod, Intrinsic, Futureway, Jet Brains, Rooflow, Haxter, Seven Sense, Google Summer of Code, Open M, our newest bronze member, Orbic, Tang Vision, AMP Software, Intuitivo, Rerun.io, the Edge AI and Vision Alliance, and Big Vision. If your company uses OpenCV and wants your name to be listed amongst these industry titans, please send one email to Phil at opencv.org and we’ll get talking about it. Um, OpenCV membership has a program designed for companies of all sizes. Whether you’re a new scrappy startup or whether you’re a huge established company like ARM or Qualcomm or whether you’re somewhere in the middle like Rooflow, OpenCV membership has a plan that will work for you and you’ll be able to help support the library that you depend on.

    Is also supported by our sponsors on GitHub. By the way, your screeners, we got 14 of them right now. What’s that? Sorry, the screen. Yeah, we’re having some local internet slowness here. It’s We’ll figure it out. Um, yes. So, uh, we’re also supported by our GitHub sponsors. You can go to github.com/sponsors/opencv

    and be as cool as Christopher. Uh, man, this is this to me looks like maybe a Polish name. I don’t know if I can do this one. Piraski. Um, we’ve got DJ. We got Tegan Burns. We’ve got Live Pier. We’ve got the homies at Rooflow, double supporters of Open CV. Jesus Anna, DJ Greenwood, IPOP.AI, we’ll have to have them back on the show sometime soon. Comet ML,

    Jonas Heinla, Alexander Ismolof, Tala Hussein, Nick Libertini, Blue J, and Ruters Laboratories. We’re also brought to you today by OpenCV’s official merch shop on at opencv.mmyspreadshop.com. We recently slashed prices on everything on the shop. It is now about 50% what it cost just a few weeks ago. For example, iconic statement um on a shirt if the page for me. We’ll we’ll see. We’ll see. $3 versus was about the the 40 that we had it set as last time. Our thinking with this this was, you know, if you want an open CV shirt, you really want for the import CV2 shirt on a very comfortable 100%. You’ve also got the open

    clothing also tote bags, pins, fridge magnets. I love OpenCV in English simplified Chinese and an espanol yoamo. Uh OpenCV. Uh you can buy these things at once again opencv.mmyspreadshop.com or you can scan the QR code that I just put up on the screen. 23 bucks for a pretty sweet Open CV t-shirt. Um, since we dropped the prizes earlier this week, we’ve actually sold uh twice as many shirts in the last week than we sold in the last year. So, uh, feeling pretty good about it. Feeling pretty good about it. Really? Yeah. Okay. Well, uh, if you if you drop the price by this much, I think we are definitely not getting any donations really. But it is for open season. Yeah. You know, I mean, it’s it’s promo, right? I think I think at this point inside baseball I think we make about five bucks a shirt or something like that. So, you know, it’s not it’s not that it’s not that much, but it it is it is cool to see people like at your CVPRs or these conferences we go to wearing OpenCV gear. So, buy a shirt. Uh you are supporting OpenCV in your in your small way. Um and we really do appreciate all of these ways for folks to donate. Um before the end of this year, we will also produce um an OpenCV 25 year limited edition t-shirt. So stay tuned for that. Yeah. Or possibly even a challenge coin if you’re into that. Um we we were just talking about this. But okay, uh it is now time for the trivia giveaway. First person to answer in the chat wins the prize. Um, somebody in on YouTube already tried to answer the question before I asked it. Um, bad news. You’re wrong. But, uh, I appreciate your gumption. I really do. Um, I think, uh, uh, you know, trying to get in my head is is is not easy. Oh, wow. Somebody just said they’re buying 10 shirts for their staff at work. Hell yeah. Oh, wow. Thank you so much. Um, thank you, Carlos. We really appreciate that. Um, so we have talked about Quen uh 2.5BL today. Uh, if you’ve won in the last couple of months, please do not answer and give someone else a chance to win. By the way, uh, if you’re answering on Zoom, everyone, and not just hosts and panelists, if you’re watching anywhere else, just post it in the chat. Um, I’ve got both of these windows up right now. We’ve talked a lot about Quen 2.5VL today. in 2.5VL release. What month and year both parts of the answer to win here today?

    You got me. Oh, you’re close, Lawrence. I think this is the first time you got me. I don’t know exact. Oh, hey, we got it. It looks like uh look looking here. I see uh Pom Pomudu Pamudu over on Zoom. uh just answered

    January. That is indeed correct was uh January 28 official Alibaba blog. Um so congratulations. I know we’re having a little bit of sorry folks. Um but the good news is the laggginess applies to everybody. Zoom is getting the same lag as YouTube. We’re getting the same lag as Twitch. Um, and so, you know, it is what it is. Um, but yes, some of you folks out there also got the right answer. Um, Pamudu is our winner for this week’s trivia giveaway. Please go to opencv.org/university.

    Pick out a course, send me an email that’s philopc.org

    with the name of the course that you would like and we will make sure you get that added to your OpenCVU dashboard. Congratulations, Pamudu. We’ve got a few questions here, Doc, while we uh before we wrap up here. Um we got somebody asks, “Can you send t-shirts overseas to Angola?” I don’t know about Angola specifically, but if you go to the opencv.mmyspreadshop.com

    and add something to your cart, it’ll have a list in the checkout of the of countries they do ship to. And I know it it’s it ships pretty fairly like almost all over the world where it’s that’s like possible to get stuff into basic basically uh without without you know grease and some palms as we as we’ve talked about on this show over the years. So uh thanks for your interest and uh hopefully we can get you a shirt. Um got a couple of questions here saved. Let me go ahead and I’m going to drop off my face so Doc can answer some of these questions himself here. Um, here’s a good one from earlier in the stream from uh from Shan. What are some common mistakes if an image is not classifying correctly and how can we fix them? Well, so uh if it is not a VLM, right? In a VLM, you can use a better prompt, right? So, you can say that let’s say uh it doesn’t happen uh but let’s say uh a simple example, right? uh let’s say it gets uh confused between uh between a dog and a wolf, right? Then uh you would run your you know the prompt you would change that pay special attention right to u dogs and wolves right so most of the things with uh you want to first use prompt engineering to get uh this image classifier if you’re using a VLM right if you’re using a VLM uh it is tricky because this is a large uh large model right you don’t want to jump into finetuning immediately uh so you want to do something using prompt engineering as much as you can. So you would give extra prompts for classes it is trying it it is getting mclassified with saying that pay special attention or you know uh make sure that you don’t do this uh do that. So it does uh those kind of prompting does affect the final output. Uh but you can also fine-tune the model uh to give better results. Right? So you can add image caption pairs to improve the quality uh of the output. But in standard image classification, right, the best way is to just fine-tune the model, right? Because those are not difficult to fine-tune. They’re they’re pretty uh easy to fine-tune. Just wherever it’s getting confused, let’s say it is getting confused between dogs and wolves. Take those examples, the hard examples, this is called hard negative mining. So you take the hard examples and put it in the training set and fine-tune, right? It’s as simple as that. find all the hard examples where the mis mclassification is happening and then uh fine-tune the model again and that pretty much fixes the problem most of the time.

    So the in if you’re using standard image classification network the mistake you have made is that your training data is not rich enough right so you basically uh fix your training data

    Phil,

    where’s Phil?

    Sorry guys, we are having some uh technical issues today because of uh uh because of the network.

    Oh, let me see if I can I can find some questions.

    All right. Uh can we okay says can we please have a session on quen image edit 2 point taken uh we’ll we’ll see we uh when we can fit that in. So we uh basically we have uh a very packed schedule for the next few weeks. So whenever we have some uh you know guest drop off or something happens then we we modify that uh session to a tutorial session. But uh that’s a that’s a uh you know I’ll I’ll look into it how much time it takes to fit that in and if it is in a one in 1 hour we can fit that in we’ll definitely try to do that. That’s a good request. Thank you.

    All right. So I’m not sure what’s uh going on. uh right now with the with the network issue. So uh but we are close to uh you know our program already and I have a hard uh stop in just a few minutes. So uh let me conclude this uh show um and and thank you Phil as I said that if anything goes wrong it is his uh it is uh you know Phil organizes this show and if anything goes wrong today it came true finally. Oh, you’re back. Um, for some reason it it put my screen up as well. I don’t know why it did that. Um, yeah. Sorry about that, folks. We we had a uh a a total internet dropout here um in uh in in Tijuana. So, I think it’s it might be because it’s raining a little bit out there. Anyway, um thank you. Uh yeah, Camille. Um maybe we will have a a Quen image edit tutorial here. I’m pulling up the chat. I think we had we had at least one more good Unfortunately, we just lost Zoom. Sorry, folks. Um, yeah, I think we’re we’re just fizzling out here. It’s okay. I think uh we can conclude today’s session and uh the questions we can answer by email.

    Son of a [ __ ] You’re Phil, you’re online.

    All right. So, uh, Phil thinks that he is offline. Uh, but we are online and, uh, so I I just want to thank everybody who is here today for this uh, for this webinar. Thank you so much. And uh also our sale is going on. If you’re interested in the OpenC CV University courses which are your guide to become an AI engineer, please uh join today. This is one of the best deals you can find. And if you purchased a course a program during this webinar, please send us an email at coursesopencv.org and we will be able to uh you’ll you’ll receive the free book. All right. Uh that’s it. Thank you so much. Uh thank you guys. U I’m uh we have to conclude this because and unfortunately Phil is not here because of network issues but we’ll see you next week.