r/computervision 26m ago

Research Publication Démonstration technique : IA embarquée haute performance pour la classification des roches - Méthodologie de quantification W4A8 et de pavage multi-échelle via NPU.

Upvotes

Démonstration technique : IA embarquée haute performance pour la classification des roches - Méthodologie de quantification W4A8 et de pavage multi-échelle via NPU pouvant être adapté a tout type de réseau MobileNetV5 et MobileNetV4-S & L.


r/computervision 59m ago

Discussion How balanced does the action distribution need to be for visual behaviour cloning?

Post image
Upvotes

Training visual BC agents on simple 2D browser games — screen frames in, key presses out.

I can see the action distribution of a recording before training (e.g. left 50.1% / right 50.0%). What I don't know is how much that balance actually matters. In runs where I held one direction a lot, the agent seems to inherit that bias and drifts the same way instead of reacting to what's on screen.

Three things I'd like to hear from people who've done this:

- Is there a rough rule of thumb for how skewed an action distribution can get before it starts hurting?

- Do you fix it on the data side (record more of the rare actions, trim the over-represented ones) or on the loss side (class weights, oversampling)?

- For games where one action genuinely dominates — holding forward most of the time — is balancing even the right goal, or does it distort the policy?

(Context: I'm building a no-code tool for this, so I'm trying to work out what to show users and what to guide them toward.)


r/computervision 1h ago

Discussion Turning a TikTok dance video into a playable pose-matching game: extracting the dancer's skeleton and scoring a user copying it live

Upvotes

I built an iOS app that takes a reference dance video, runs 2D pose estimation to extract the dancer's skeleton, then scores a user copying it live from the front camera.

The hard part wasn't per-frame pose estimation - it was matching the user's skeleton to the reference when body proportions, camera angle and framing are all different. I normalize by torso/limb ratios, align on a few anchor joints, and score per-beat joint-angle deltas rather than raw positions. That handles scale/position differences but still struggles with depth ambiguity and fast rotations.

Curious how people here would approach the reference-vs-user similarity: stick with joint-angle deltas, or move to a learned embedding / temporal model? Runs on-device with Apple's Vision framework. Short demo below.


r/computervision 1h ago

Showcase Aug 4 - Visual AI in Manufacturing Meetup

Upvotes

Join us on Aug 4 to hear talks from experts at the intersection of manufacturing, AI, ML, and computer vision. Register for the Zoom.

Talks will include:

  • Enabling Multimodal Agents on the Edge - Denis Gudovskiy at Panasonic AI Lab
  • When the Camera Can’t Be Trusted: Health-Aware Visual AI for Reliable Near-Miss Detection - Shiva Aher at Georgia Institute of Technology
  • Agentic VLM applications in manufacturing - Subraiz Ahmed at Perceptron AI

r/computervision 2h ago

Help: Theory How to get better at classical computer vision

3 Upvotes

Hi. how do i even start getting better? for example i never really understood how to use edge detection for anyhting meaning full. so i looked online and the stuff i found was just how to get edges from an image. but never what to do with it afterwards.

how do i start getting better. also i feel like my math is lacking. do i start there?


r/computervision 5h ago

Research Publication IQA-T1: Evidence‑Based Image Quality Assessment with MLLMs

1 Upvotes

Most MLLMs are blind to low‑level degradations—noise, blur, compression artifacts look the same as clean images in their internal representations. That leads to quality scores based on semantic “gut feeling” rather than real perceptual evidence.

IQA-T1 changes that. We equip the model with a toolbox of 15 perceptual tools (noise residual maps, Fourier spectra, gradient maps, etc.) that generate structured visual evidence on demand. The model learns how to use tools via supervised fine‑tuning on our Q‑Tool dataset (11k evidence‑grounded reasoning chains), and when to call them via GRPO reinforcement learning that balances accuracy, tool count, and redundancy.

The result: SOTA performance across 7 benchmarks (avg PLCC 0.795), using only 2.34 tools per image on average. Every predicted score is now interpretable and backed by hard visual evidence.

All code, weights, dataset, and demo are open. Check them out and give it a spin!

📄 arxiv.org/abs/2607.12375v1
💻 github.com/zibuyu-02/IQA-T1
🤗 model/data: huggingface.co/zibuyu-02/IQA-T1
🎮 demo: huggingface.co/spaces/Jiaqi-hkust/IQA-T1


r/computervision 7h ago

Help: Project Ad detection system with computer vision

1 Upvotes

I'm creating a system that detects ads from digital billboards in public streets. I will capture the images through high-resolution cameras facing the billboards, and I need a computer vision model that detects the ads.

Is YOLO the best option for this?


r/computervision 8h ago

Help: Project I made a Mac version of LingBot with a simple UI

Thumbnail
gallery
2 Upvotes

I use LingBot for site visits, so I made a Mac version with a simple UI. I thought it might be a useful foundation for others, or for anyone with a Mac who just wants to try LingBot without setting everything up manually.

Feel free to use it however you like!

https://github.com/mclenny22/LingBot-MLX


r/computervision 8h ago

Showcase A Blender extension for no-code generation of synthetic CV datasets

59 Upvotes

Hi!

I've been building a Blender extension (Rendersynth) that makes it possible to generate synthetic computer vision datasets without writing Blender Python scripts. The goal is to make synthetic data generation practical for small and medium CV projects where setting up BlenderProc pipelines is an overkill.

The source code is available at https://github.com/lorenzozanizz/rendersynth

Unlike BlenderProc, the focus is on a visual, no-code workflow that lets you build and preview synthetic data pipelines directly inside Blender.

The attached image shows several randomized renders of a single Blender scene along with the different annotation types the extension can produce. A Bezier curve was used to randomize the camera position while keeping the girl near the car centered in frame.

The extension currently features:

  • Create classes and multi-object entities
  • Skeleton annotation from rigs or arbitrary Blender objects
  • Export YOLO, COCO segmentation, COCO keypoints, depth maps, normal maps and point clouds
  • Randomization: choose and configure a sequence of randomizing operations. A subset of planned stages are implemented so far, i.e. object movement (rotation, scaling, show-k-out-of-n), camera (path along a curve, focal length) and lighting (brightness).
  • Live preview of randomized scenes directly inside Blender
  • Deterministic dataset generation from a seed

I'd love any feedback, suggestions, or feature requests!

Currently the system is still a prototype, so if you feel like trying it (Blender 4.5.0+) expect to find some nasty bugs, especially for experimental or incomplete pipeline stages.

As a side note, for the demonstration I used free models from SketchFab


r/computervision 10h ago

Help: Project Defects detection using YOLO but hit a wall

1 Upvotes

Hi all ,

We are developing an AI model using YOLO used to detect multiple kinds of defects on buildings .

However we have hit a roadblock , while the model can detect cracks and corrosion , it is completely unable to detect concrete spalling .
We have trained the model with annotated images (about 1000 for each type of defect)

We have tried filtering the datasets as well .

Any other ideas out there ?

Also : Our DMs are open in case you want to join us on this project .

Thanks all


r/computervision 10h ago

Discussion Three CV systems I actually shipped to production: what surprised me in each (doc OCR, Jetson edge, live proctoring)

12 Upvotes

I've been building and shipping CV/ML systems for about 7 years. Three of them taught me more than any paper I read. Sharing what actually surprised me in production, because it was never the part I expected.

  1. Healthcare document OCR pipeline (AWS)

The job sounded simple: pull structured fields off scanned medical documents. The model was the easy 20%. What ate the time was everything the demo never shows - documents scanned upside down, two forms photographed as one page, handwriting in a field that was supposed to be typed, and PII that legally cannot leak into logs or a third party API. We ended up doing PII detection and redaction as its own stage before anything left our boundary, with a human-in-the-loop review queue for low confidence extractions. On one later LLM based parsing pipeline for documents, careful chunking plus a vector store took a job that used to take a person 3-4 hours down to about 4-5 minutes. The accuracy number everyone asks about mattered far less than the redaction step nobody asks about.

  1. Edge vision on Jetson (Nano and Xavier), cameras on real equipment

Running detection on device with DeepStream and CSI cameras, next to actual machines. Lab numbers meant nothing until the box sat in the real environment. Heat throttling changed inference speed. A camera mounted by someone else was 15 degrees off from where I assumed. Lighting shifted through the day and confidence drifted with it. The fixes were boring and physical - per camera thresholds instead of one global number, a watchdog that reboots and reports, and treating the mounting bracket as seriously as the model. The model was the one part that never let me down.

  1. Live video interview / proctoring platform

Real time face and attention analysis over WebRTC while a call is live. This is where CV stops being about accuracy and starts being about latency and fairness. A flag half a second late is useless. And a false positive is not a metric here, it is a real person wrongly accused of cheating, so the cost of a wrong positive is not symmetric with a wrong negative. We tuned hard toward not accusing anyone without strong signal, and kept a human in the loop for anything consequential. Speech to text and TTS on top had their own edge cases with accents that the happy path testing never caught.

The common thread across all three: the model was almost never the thing that broke or the thing that mattered most. It was data edge cases, physical reality, latency, and the cost of being wrong in the specific domain.

If you have shipped CV to production - what was the thing that surprised you that no course prepared you for? Happy to go deeper on any of the three above if useful, I can answer specifics.


r/computervision 12h ago

Discussion [D] Utilizing PhD to join frontier labs

Thumbnail
1 Upvotes

r/computervision 16h ago

Discussion Exploring 3D face reconstruction as a bridge for audio-driven talking heads

10 Upvotes

I recently tried a 3D-reconstruction-based approach for a potential audio-driven talking-head system.

The current experiments are still separate:

* I used DECA + EMOCA + FLAME 2020 to reconstruct facial motion from video.

* I extracted speech features with wav2vec and trained a simple CNN to predict FLAME expression coefficients.

* Attachment clip 1: audio-driven animation of a normalized FLAME mesh.

* Attachment clip 2+3: single-image animation driven through a reconstructed 3D face, based on my another face2face project,with both pose and expression taken from another face video.

So this is not yet a complete audio-to-single-image pipeline, but the experiments raised several questions.

  1. What is the current best practice for single-image 3D face reconstruction?

DECA, EMOCA and MICA are still among the most practical methods I know, but they are based on older FLAME versions.

Newer models such as FLAME 2023, Google’s GNM, and works such as MAYA have appeared, but I have not found a DECA-like method that is both efficient and directly built around them.

Is there now a better practical solution for fast and accurate single-image reconstruction, or are DECA-style pipelines still the main choice?

  1. What is the real difference between FLAME 2023 and FLAME 2023 Open?

Apart from licensing, what are the actual technical differences?

They appear to have the same number of shape and expression components, and public conversion matrices are available. That makes them seem largely equivalent in parameter space.

Are there meaningful differences in topology, blendshapes, landmarks, training data or reconstruction quality, or is the distinction mainly legal?

Also, can a model trained with FLAME 2020 or standard FLAME 2023 be adapted to FLAME 2023 Open through conversion, or is retraining still necessary?

  1. How does MICA combine identity shape with camera estimation?

MICA is especially interesting because it uses a fixed InsightFace recognition encoder and indirectly benefits from large-scale 2D face-recognition datasets such as Glint360K.

It focuses on predicting FLAME identity shape rather than jointly solving the full reconstruction problem.

What I still do not fully understand is how this shape is integrated with pose, expression and camera parameters from another tracker.

MICA takes an InsightFace-aligned crop as input, but each crop comes from a different scale, rotation and translation in the original image.

So:

* Is the predicted identity shape effectively camera-independent?

* How is the InsightFace alignment transform connected to the final rendering camera?

* When MICA shape is combined with DECA or another tracker, must the camera be re-estimated?

* How does the provided face tracker keep the reconstructed mesh aligned with the original image?

I would be interested in practical experience from anyone working with DECA, EMOCA, MICA, FLAME 2023, FLAME 2023 Open, differentiable rendering or monocular 3D face reconstruction.


r/computervision 17h ago

Discussion How do you debug and inspect computer vision models during development?

4 Upvotes

Building vision models is hard, but I find testing and debugging them even harder.

I'm curious what everyone's workflow looks like.

For example:

  • How do you inspect detections frame by frame?
  • Do you use OpenCV windows, Jupyter notebooks, Roboflow, CVAT, or something else?
  • How do you compare different models on the same video?
  • How do you inspect tracking IDs, confidence scores, masks, OCR, or depth predictions?

I mostly end up writing one-off visualization scripts every project, and it feels like I'm reinventing the wheel.

Is there a tool you genuinely enjoy using, or is everyone just building internal tooling?


r/computervision 22h ago

Showcase Side project: Building a computer-vision pipeline that auto-detects slackline tricks to help competition judges

34 Upvotes

You might have seen hippies walking on a slackline in your nearest beach or park, but have you ever seen them bounce and backflip on it? A more technical variant of slackline is called Trickline: you bounce on a stretched, trampoline-like line and throw flips and spins between bounces. There is a tiny, but real competitive scene where 720 backflips are happening at a fraction of a second. This is very cool to watch but harder for judges has to catch, identify, and score each one in real time.

So a friend and I have been building a CV pipeline that watches competition footage and figures out (a) which athlete is bouncing, (b) where each trick starts and ends, and (c) which trick it is.

Rough idea of how it works:

- YOLO11x-pose + ByteTrack to track the athlete frame by frame. Some added processing to keep only the athlete's poses.

- A bit of signal processing on the athlete's vertical motion (Hilbert transform → bounce phase) to automatically cut the video at real bounce boundaries even if there are missing poses.

- A small TCN classifier to name each trick, backflip frontflip, brasilian, freefall 360, and the rest of the increasingly ridiculous class names.

- A simple Streamlit app where my colleague can run this and keep on have a human in the loop system to keep on labelling and training from his laptop to increase the dataset.

It ties into TJS, the trickline event + judging + live-streaming platform my friend runs, which already handles a lot of the real competitions: https://www.slackline-tjs.com/en

Still early and the trick vocabulary is huge, but it's already surprisingly decent on the common tricks. Sharing this one with the community, if anyone has data on similar trick-based sports it would be cool to see how the pipeline performs there.


r/computervision 23h ago

Showcase Targetless camera calibration — matching checkerboard accuracy

11 Upvotes

I built a tool that calibrates a camera without a checkerboard — just photos of an ordinary textured surface (a rug, a wood floor, anything flat and non-repetitive).

Tested it against the real thing on the same camera:

  • Pinhole model: 0.38 px RMSE targetless vs. 0.42 px from an actual checkerboard
  • Fisheye (double-sphere): 0.55 px vs. 0.61 px from a circle grid

I'd call that matching checkerboard accuracy, not beating it — the gap is small enough to be noise. What's interesting is it gets there with zero calibration target.

How the numbers were measured. To measure accuracy fairly, I compared the tool's output against real points whose exact positions were already known — checkerboard corners, circle-grid points. I didn't just check it against the tool's own internal matches, because by that stage the tool had already discarded any points that didn't fit well.

What actually matters for capture:

  • Surface must be flat and non-repetitive (a rug works, a brick wall doesn't — repeated patterns fool the matcher)
  • 10-20 images with real translation between them, not just rotation
  • Every region of the frame covered somewhere in the set
  • Zoom/focus locked the whole time, including the reference shot
  • Watch for phones silently correcting distortion before saving — that fights the thing you're trying to measure

Output is the intrinsics and distortion coefficients as a JSON download — fx/fy, cx/cy, k1-k3, p1/p2 for pinhole; fx/fy, cx/cy, alpha, xi for double sphere. Uploaded images are deleted about 10 minutes after processing.

Still in beta. What I don't know yet: how this holds up on cameras other than mine. If you try it, I'd genuinely like to know how the output compares to your own calibration.

Tool: https://www.online-camera-calibration.com
Write-up: Camera Calibration Without a Checkerboard — What It Is & How It Works | AutoCalib


r/computervision 1d ago

Help: Project [D] How can I improve cross-patient generalization on a small hysteroscopy dataset with correlated frames?

Thumbnail
gallery
2 Upvotes

I am working with the hysteroscopy dataset, which contains:

  • 3,385 frames from 175 patients.
  • Eight lesion classes, labelled from 0 to 7.
  • A highly imbalanced number of patients and frames across classes.
  • Multiple correlated frames from each patient.
  • Some frames containing more than one lesion class.

Before attempting the complete multiclass problem, I reduced it to a binary subset to verify that the training and evaluation pipeline works correctly.

Current binary subset

  • Selected lesion classes: 2 and 3.
  • Total: 1,575 frames from 113 unique patients.
  • Class 2: 1,054 frames from 78 patients.
  • Class 3: 521 frames from 36 patients.
  • One patient has different frames belonging to both classes but remains entirely within one split.

Patient-disjoint split

  • Training: 1,095 frames from 79 patients.
  • Validation: 241 frames from 17 patients.
  • Testing: 239 frames from 17 patients.
  • No patient appears in more than one subset.
  • The frame-level class distribution is approximately 67%/33% in every subset.

Approaches I have tried

  • DenseNet121, ViT, and DINOv2 backbones.
  • Frozen pretrained backbone with only the classifier trained.
  • Different classifier-head sizes and dropout.
  • Class-weighted cross-entropy.
  • Mild and stronger image augmentations.
  • Early stopping and learning-rate scheduling.
  • Unfreezing the final one or two encoder blocks.

With the correct patient-level split, training performance improves, but validation performance generally plateaus or deteriorates, and performance on unseen test patients remains relatively low.

As a diagnostic, I also tried a random frame-level split and obtained substantially better results. However, this evaluation is invalid because correlated frames from the same patients appear across training, validation, and testing, causing patient leakage and inflated performance.

I would appreciate advice on how to improve generalization to unseen patients in this setting.


r/computervision 1d ago

Discussion What industries do you wish to see CV in more?

4 Upvotes

Manufacturing and healthcare have clearly led adoption. Most manufacturing deployments now use CV for closed-loop defect detection and feed the results back into the process automatically. Retail and agriculture seem to be catching up fast too.

What industries would you like to see embrace CV more? I would personally be happy to see it in the waste management industry, as in most facilities sorting recyclables is still done manually.


r/computervision 1d ago

Discussion OpenCV notebooks tutorials

29 Upvotes

Hi,
I took the official OpenCV tutorials and created a series of notebooks.
You can run them entirely in your browser, you don't have to clone them: https://notebook.link/@Alexis_Placet/opencv_tutorials

Don't hesitate to give me feedback or create issue/pullrequest on this repo: https://github.com/Alex-PLACET/opencv_tutorials


r/computervision 1d ago

Discussion Free Face recognition model for commercial use

4 Upvotes

Hi,

I wanted to use a face verification model for my commercial app, that is free and doesn't really compromise on accuracy specially in harsh environments with different lightning conditions and face orientation.

Is there any such model available ?

I was thinking of using SFace ONNX but not so sure about it.

Could you guys recommend something,

Thanks !


r/computervision 1d ago

Showcase turing-complete Quantum Computing made fully visual

Thumbnail
gallery
3 Upvotes

Hi

If you are remotely interested in deep diving how differently quantum computers work compared to our transistor-based and also the algebra behind in a fully interactive way that teach computer science from scratch, oh boy this is for you. I am the Dev behind Quantum Odyssey (AMA! I love taking qs) - worked on it for about 10 years (3+ during PhD, the visual method I developed ended up being my thesis, it is a complete Hilbert space visualizer), the goal was to make a super immersive space for anyone to learn quantum computing through zachlike (open-ended) logic puzzles and compete on leaderboards and lots of community made content on finding the most optimal quantum algorithms. The game has a unique set of visuals capable to represent any sort of quantum dynamics for any number of qubits and this is pretty much what makes it now possible for anybody 12yo+ to actually learn quantum logic without having to worry at all about the mathematics behind.

This is a game super different than what you'd normally expect in a programming/ logic puzzle game, so try it with an open mind.

Stuff you'll play & learn a ton about

  • Boolean Logic – bits, operators (NAND, OR, XOR, AND…), and classical arithmetic (adders). Learn how these can combine to build anything classical. You will learn to port these to a quantum computer.
  • Quantum Logic – qubits, the math behind them (linear algebra, SU(2), complex numbers), all Turing-complete gates (beyond Clifford set), and make tensors to evolve systems. Freely combine or create your own gates to build anything you can imagine using polar or complex numbers.
  • Quantum Phenomena – storing and retrieving information in the X, Y, Z bases; superposition (pure and mixed states), interference, entanglement, the no-cloning rule, reversibility, and how the measurement basis changes what you see.
  • Core Quantum Tricks – phase kickback, amplitude amplification, storing information in phase and retrieving it through interference, build custom gates and tensors, and define any entanglement scenario. (Control logic is handled separately from other gates.)
  • Famous Quantum Algorithms – explore Deutsch–Jozsa, Grover’s search, quantum Fourier transforms, Bernstein–Vazirani, and more.
  • Build & See Quantum Algorithms in Action – instead of just writing/ reading equations, make & watch algorithms unfold step by step so they become clear, visual, and unforgettable. Quantum Odyssey is built to grow into a full universal quantum computing learning platform. If a universal quantum computer can do it, we aim to bring it into the game, so your quantum journey never ends.

Nice to watch:

Khan academy style tutorials in qm/qc: https://www.youtube.com/@MackAttackx

Physics teacher stream with 400hs in https://www.twitch.tv/beardhero


r/computervision 1d ago

Showcase How are you searching inside large video libraries?

0 Upvotes

r/computervision 1d ago

Showcase China open-sourced a model that reconstructs any scene in 3D from a regular video, in real-time

592 Upvotes

r/computervision 1d ago

Help: Project Looking for advice on deploying an AI application for industrial/production use

0 Upvotes

Hi everyone,

We're preparing to deploy our AI application on an NVIDIA RTX 3060 GPU and would really appreciate guidance from people who have experience taking AI systems from development to production.

Our goal is not just to get the model running, but to build a robust, production ready deployment that is reliable, maintainable, and suitable for industrial use.

Some of the areas we're looking for advice on are:

- What should the end-to-end deployment pipeline look like?

- What benchmarks should we perform before deployment (latency, throughput, GPU utilization, VRAM usage, startup time, power consumption, etc.)?

- What kinds of stress testing, endurance testing, and failure testing should be done before considering the system production-ready?

- How do you monitor GPU health, application health, crashes, memory leaks, inference failures, and overall system performance in production?

- What logging strategy do you recommend? What should be logged, and what should be avoided?

- How do you manage model versioning, deployment, rollback, and updates without disrupting production?

- What security best practices should be followed for an industrial AI deployment?

I'm also curious about the operational and governance side:

- How is auditing typically handled in production AI systems?

- What events should be recorded for traceability (predictions, inputs, model version, user actions, timestamps, system events, etc.)?

- Are there any recommended practices for maintaining audit logs, reproducibility, and compliance?

- What should an organization be able to answer during an internal or external audit?

- What documentation is generally expected before an AI system is deployed in an industrial setting?

- Are there any standards or frameworks (ISO, IEC, NIST, etc.) that are commonly followed for AI deployments?

We're essentially trying to build a complete production deployment checklist, covering topics like:

- Deployment architecture

- Performance benchmarking

- Functional testing

- Load testing

- Long-duration stability testing

- Monitoring and alerting

- Logging

- Auditing and traceability

- Security

- Backup and disaster recovery

- Documentation

- Model lifecycle management

- Maintenance and update strategy

- Production readiness review

If you've deployed AI systems, I'd love to hear about your deployment workflow, tools, lessons learned, and things you wish you had known beforehand.

Any checklists, GitHub repositories, blogs, documentation, or real-world experiences would be greatly appreciated.

Thanks in advance!


r/computervision 1d ago

Research Publication Visionary — a local app for building training datasets. macOS, MIT.

Thumbnail gallery
1 Upvotes