Method

A data program is a system. Each stage is a layer that locks into the one before it, so the result can be trusted from the first recording to the final delivery.

01Human experience

Design the data program

We start with what the model must understand, then define the environments, activities, coverage, and acceptance criteria that will get it there.

02Observation

Capture real-world signals

Trained contributors record activity in real settings, using the modalities the task calls for: video, audio, sensor, spatial, and written context.

03Context

Structure and annotate

Raw material becomes organized, labeled data: boundaries, keypoints, sequences, transcripts, and the metadata that keeps it meaningful.

04Structure

Validate for quality and relevance

Multi-stage review checks accuracy and consistency, and confirms the data still answers the original question.

05Intelligence

Deliver data ready for AI development

Documented, versioned, and traceable datasets, delivered in the format your team builds with.

How programs run

The specifics buyers usually ask for first.

01

Capture rigs and modalities

  • Head-mounted egocentric cameras for first-person task data
  • Fixed multi-view cameras for spatial and multi-person context
  • Wearable and environmental sensors where the task calls for them
  • Close-talk and multi-party audio with dialect-aware transcription
  • Written context: checklists, forms, screens, and instructions
02

Delivery formats

  • MP4 and WAV with JSON or JSONL sidecars
  • COCO and YOLO for detection and keypoints
  • RLDS and LeRobot-compatible episodes for robotics
  • WebDataset shards and Parquet tables for training pipelines
  • Your schema, if you already have one
03

Quality process

  • Acceptance criteria written and signed before capture
  • Three review passes: self-check, peer review, program lead audit
  • Inter-annotator agreement measured on an overlap set for every label type
  • Coverage tracking against the program plan, so gaps are visible early
  • Every data point carries origin, reviewer, and acceptance status

What a pilot looks like

Four weeks from first call to delivered data. Sized to prove quality against your own benchmarks before anyone commits to scale.

01

Week 1: design

Scope the task and environments, agree acceptance criteria, brief contributors, confirm consent and governance.

02

Weeks 2 to 3: capture and structure

Typically 20 to 60 hours of validated data across two or three sites, structured as it arrives.

03

Week 4: review and delivery

Final audit, documentation, and delivery in your format. A written read-out on what worked and what should change at scale.

How we work together

Three shapes of engagement. Every one starts with written acceptance criteria and ends with a documented delivery.

01

Pilot

Four weeks, fixed scope, fixed price. Enough data to evaluate quality against your own benchmarks.

02

Program

Eight to sixteen weeks, priced per hour of validated data or per program, with agreed coverage targets.

03

Ongoing

A standing program with a named team, monthly deliveries, and versioned datasets.

What you receive

  • A written program design with scope, coverage, and acceptance criteria
  • Documented consent and governance for the program
  • Structured, validated data in your schema
  • Provenance and review records for every delivery
  • A named team that stays with the program

What we need from you

  • A clear statement of what the model must understand
  • The environments and tasks it will work in
  • Your data formats, schemas, and any existing guidelines
  • A point of contact for decisions on scope and acceptance

Tell us what your AI needs to understand. We'll help shape the data program that gets it there.

Preview · example content