The Challenge
The client approached us with what appeared to be a straightforward data collection request: short videos of cats and dogs moving indoors. This was a highly controlled data collection initiative designed to train and refine a computer vision model focused animal motion understanding.
The model required clean, consistent footage of the animals indoors with clear guidelines to ensure usability for machine learning. Each video had to be 10-30 seconds max, filmed indoors, with a single pet, fully visible in frame and moving. The camera had to remain steady to avoid introducing motion noise that could interfere with downstream pose estimation and gait analysis.
The final dataset needed diversity across species, breeds, sizes, and age groups. Metadata capture was equally important, including device type and structured categorization attributes to support supervised learning.
In short: this wasn’t just cute pet footage. It was controlled training data for a computer vision AI system.
The Solution
We began with a small managed collection pilot of 50 videos using internal resources to validate the project and assess feasibility. Once validated, we scaled collection through our global contributor network DATAtalent who worked within AIDA Hub, DATAmundi’s central platform designed to handle the complexities of multimodal data collection and management for AI projects.
Our AI contributors were onboarded with the guidelines and using AIDA Hub, uploaded up to five clips per pet, and also submitted structured metadata including breed, size category, age group, and recording device.
Some submissions surfaced practical edge cases that matter in machine learning training. For example, contributors initially used food to get the animal to move. This caused pets to lower their heads too close to the floor. This obscured body posture and you couldn’t see the skeletal and movement clarity.
ADIA Hub enforced technical constraints automatically, including video length verification and duplicate detection. The system flagged repeated uploads – whether the same video was submitted twice or identical footage appeared from different contributors – verifying contributors and protecting dataset integrity.
Every video then passed through manual human QA from DATAmundi’s specialist quality team. Each clip was reviewed to confirm compliance with movement requirements, visibility standards, and camera stability. Where needed, contributors were directed to reshoot, maintaining strict quality thresholds without disrupting delivery cadence.
We delivered approximately 200 validated videos per week over a two-month period, incorporating continuous client feedback.
The Result
Within two months, we collected and delivered 1,000+ approved, high-quality video samples suitable for supervised specialist model training that aligned with the client’s specialist guidelines.
All duplicate content was eliminated, technical specifications were enforced, and delivery was on schedule. In several cases, contributors were so keen and enjoying the project so much, one participant who sourced footage from 50 different pets by recruiting friends and family to the project!
The result was not just volume, but a controlled, structured dataset ready for ingestion into a computer vision training pipeline.
The Business Impact
This project helped the client train a computer vision model capable of understanding how animals move in real indoor environments.
It also proved something we know well: seemingly simple data collection tasks become complex very quickly when quality truly matters. They say you should never work with animals or children, but this project proved to be a great success.
By combining a robust data management orchestration platform, good workflows, and human expertise, we were able to turn unpredictable subjects into reliable training data.
Complicated? Absolutely.
But also, undeniably great fun for everyone involved.
If you’re looking to collect high-quality data at scale, we’d love to help. Connect with us here and let’s explore what’s possible.


