Method
A data program is a system. Each stage is a layer that locks into the one before it, so the result can be trusted from the first recording to the final delivery.
Design the data program
We start with what the model must understand, then define the environments, activities, coverage, and acceptance criteria that will get it there.
Capture real-world signals
Trained contributors record activity in real settings, using the modalities the task calls for: video, audio, sensor, spatial, and written context.
Structure and annotate
Raw material becomes organized, labeled data: boundaries, keypoints, sequences, transcripts, and the metadata that keeps it meaningful.
Validate for quality and relevance
Multi-stage review checks accuracy and consistency, and confirms the data still answers the original question.
Deliver data ready for AI development
Documented, versioned, and traceable datasets, delivered in the format your team builds with.
How programs run
The specifics buyers usually ask for first.
Capture rigs and modalities
- Head-mounted egocentric cameras for first-person task data
- Fixed multi-view cameras for spatial and multi-person context
- Wearable and environmental sensors where the task calls for them
- Close-talk and multi-party audio with dialect-aware transcription
- Written context: checklists, forms, screens, and instructions
Delivery formats
- MP4 and WAV with JSON or JSONL sidecars
- COCO and YOLO for detection and keypoints
- RLDS and LeRobot-compatible episodes for robotics
- WebDataset shards and Parquet tables for training pipelines
- Your schema, if you already have one
Quality process
- Acceptance criteria written and signed before capture
- Three review passes: self-check, peer review, program lead audit
- Inter-annotator agreement measured on an overlap set for every label type
- Coverage tracking against the program plan, so gaps are visible early
- Every data point carries origin, reviewer, and acceptance status
What a pilot looks like
Four weeks from first call to delivered data. Sized to prove quality against your own benchmarks before anyone commits to scale.
Week 1: design
Scope the task and environments, agree acceptance criteria, brief contributors, confirm consent and governance.
Weeks 2 to 3: capture and structure
Typically 20 to 60 hours of validated data across two or three sites, structured as it arrives.
Week 4: review and delivery
Final audit, documentation, and delivery in your format. A written read-out on what worked and what should change at scale.
How we work together
Three shapes of engagement. Every one starts with written acceptance criteria and ends with a documented delivery.
Pilot
Four weeks, fixed scope, fixed price. Enough data to evaluate quality against your own benchmarks.
Program
Eight to sixteen weeks, priced per hour of validated data or per program, with agreed coverage targets.
Ongoing
A standing program with a named team, monthly deliveries, and versioned datasets.
What you receive
- A written program design with scope, coverage, and acceptance criteria
- Documented consent and governance for the program
- Structured, validated data in your schema
- Provenance and review records for every delivery
- A named team that stays with the program
What we need from you
- A clear statement of what the model must understand
- The environments and tasks it will work in
- Your data formats, schemas, and any existing guidelines
- A point of contact for decisions on scope and acceptance
Tell us what your AI needs to understand. We'll help shape the data program that gets it there.