Reka Responsible AI, Model Risk, Ethics & Governance Framework

RekaDaily-10k: Collecting 10,000+ Hours of Egocentric Household Manipulation Data

RekaDaily-10k: Collecting 10,000+ Hours of Egocentric Household Manipulation Data

To understand and simulate the physical world, omni world models and vision-language-action models need more than text. They need high-quality visual data with descriptions of what is happening, recorded directly in the chaotic environments where real behavior occurs.

Search for cooking videos and you get an enormous library of edited, staged, tripod-mounted footage cut to keep the interesting parts. What a machine needs in order to learn a physical task is the opposite: one continuous first-person view of somebody actually doing it, at the speed they actually do it, in the mess they actually live in. Nobody uploads that, because nobody would watch it.

That footage has to be commissioned. Teleoperated data is precise but slow to produce, and it tends to inherit the tidiness of the space it was recorded in. Synthetic scenes scale but smooth over the clutter of real homes. Real first-person recordings sit in between, and the reason there are not more of them is that somebody has to pay people to make them.

Introducing RekaDaily-10k

To meet our own requirements for scale and quality, and to satisfy the needs of the broader industrial community, we built Claru, Reka’s foundational data engine, which sources egocentric video through a global network of paid collectors.

Today we are releasing RekaDaily-10k: unscripted, first-person recordings of everyday household life, captured across homes by paid collectors recording their own routines, with a significant share in 4K. We are contributing this dataset to the research community as part of our open ecosystem initiative, under Apache 2.0, and the raw tier is available now on Hugging Face, with the full 10,312 hours live by early next week.

The dataset ships in two tiers:

  • Raw tier. Unfiltered, raw footage, allowing teams to implement their own clipping, filtering, and annotation workflows. It has 10,312 hours.

  • Processed and captioned tier. Footage that has been through our processing pipeline, cut into shorter clips and captioned, for teams that want language supervision out of the box.

The Apache 2.0 licence covers commercial use and redistribution. Each clip ships as video with a text caption.

The same recordings, released twice — pick where you want to start

10,000 HOURS AS RECORDED

Unscripted first-person household sessions from paid collectors

RAW TIER

The footage as collected

hours

as recorded

segments

full sessions

captions

none

best for

teams with their own pipeline

PROCESSED & CAPTIONED

Through the pipeline

hours

over 10,000+

segments

short clips

captions

one per clip

best for

language supervision out of the box

The release is Apache 2.0 and ungated, covering commercial use and redistribution. Roughly 1,670 hours are native 4K.

The Egocentric Landscape

The egocentric ecosystem has grown quickly, and this release is meant to add to it with more videos showing egocentric activities. Ego4D established the modality and remains the reference corpus for daily life. Egocentric-10K and the larger releases that followed it showed how far first-person data can scale, and they cover industrial work in real production environments. EPIC-KITCHENS is still the gold-standard of ego-centric benchmarks of human activities.

What RekaDaily-10k adds: Unscripted household activity recorded in real homes with detailed descriptions. Roughly 1,670 hours are in native 4K, which is a higher resolution than most large egocentric corpora carriers. Together, this makes it an ideal complement for teams building domestic AI where language supervision matters.

Where The Footage Comes From

Claru is a paid collection network of more than 100,000 people recording the physical world. Collectors join a project, pass a qualification assessment, then record, submit, and get paid per accepted hour. The work spans domestic life, commercial environments and skilled trades across several regions. This release draws on the household portion of that network, recorded by a subset of those collectors on phones in head mounts.

Environment Diversity

Because every collector records in their own home, the number of distinct environments rises with the number of people who contribute rather than with the number of hours recorded. Different kitchens, appliance models, cabinet layouts, floor plans, lighting conditions and degrees of clutter. Different outlets and switch plates, different packaging on the shelves, different signage languages, different weather through the windows. That kind of variation is difficult to produce any other way.

Authentic and Unscripted

The footage is unscripted, capturing the dirt and noise of the real world, and thus making the dataset valuable. Collectors record real activity instead of performing a task list, so sessions run long, hands leave frame, tasks get abandoned and resumed later, and people walk in and interrupt. That is roughly what a deployment environment looks like, and it is the part that is hard to arrange deliberately.

Everyday Household Routines

What people recorded is mostly household work and the routine around it. Laundry from the pile to the folded stack. Kitchen cleanup, dishes, unloading groceries, wiping surfaces. Reorganising rooms, closets and drawers. Sweeping, taking out the trash, watering plants, clearing a table, unboxing something new, changing a bulb or a battery. The short fiddly two-handed jobs are here precisely because no one would commission a teleop session for them, and they are still on everybody’s actual list.

Consent and Privacy

Paid Collector Consent

This is footage recorded inside people’s homes, and publishing it openly raises the stakes on how it was obtained. Collectors are paid contractors who opt in, and every session is recorded with the wearer’s knowledge and agreement.

Bystanders

The wearer is not the only person a home camera sees, so bystanders are handled separately. Collectors are instructed to record only with the agreement of other adults present and to keep others out of frame where that is not possible, and footage that shows identifiable non-participants is flagged for review before it can be released.

PII Screening

Personally identifiable information cannot be reduced to a yes or no question. Recording real environments at this volume will inevitably lead to PII passing in front of the lens; what matters is identifying it and deciding what to do.

Three gates stand between a recording and the public release, and one of them branches

01

Collector consent

Paid contractors who opt in. Every session is recorded with the wearer’s knowledge and agreement.

02

Bystanders

Record only with the agreement of other adults present, and keep others out of frame where that is not possible.

03

Automated PII screen

Reads every clip for the cases that recur: a face in a mirror or on a screen, a name on a prescription label, a piece of mail.

FLAGGED

A human reviewer decides

— anonymise the region

— or drop the clip

CLEAR

Goes into the release

— nothing the screen and the reviewer have not cleared

Screening is not perfect. If you find something in the release that should not be there, tell us and we will remove it.

From Submissions to Captioned Video

Ten thousand hours of unedited phone footage is the raw tier. The processed tier runs quality controls, cuts the video into shorter clips, deduplicates shots, and captions them. All done at a large scale.

Quality control decides what reaches the processed tier — and it is tuned to keep, not to cut

01

Submission

an accepted hour of footage

02

Quality control

two passes, then a verdict

03

Clip & caption

deduplicated, then captioned

04

Processed tier

language supervision included

The raw tier ships straight from 01 — none of this is applied to it.

QC PASS · SIGNAL

— hands never visible

— illumination too poor to read

— frozen or duplicated frames

— wrong orientation

— solid-colour padding

— clips too short to use

— does not match the assigned activity

QC PASS · CONTENT

— not point-of-view

— staged or screen-recorded

— time inflation to pad a payout

THE RULE

— pass — straight through

— borderline — sent to review, never binned

— reject — only above a confidence floor

The thresholds were replayed against a large sample of submissions that already carried a human verdict, and moved toward the reviewers wherever the two disagreed. A false rejection is the worst outcome.

Quality control runs first, on every submission. It was introduced partway through collection, so it covers most but not all of the dataset we are releasing now; everything going forward passes through it. The checks cover the failure modes that actually recur: hands never visible, illumination too poor to make anything out, frozen or duplicated frames, wrong orientation, solid-colour padding, clips too short to use, footage that does not match the assigned activity. A second pass reads the content and asks whether it is what it claims to be, flagging non-POV, staged or screen-recorded video, and time inflation where somebody stalls to pad a per-hour payout. Every video also gets a perceptual fingerprint, because at this volume near-duplicates stop being hypothetical.

The quality thresholds came from data rather than instinct. We took a large sample of historical submissions that already carried a final human verdict and replayed them through the pipeline to compare machine flags against what reviewers had decided. Where the two disagreed we moved the thresholds toward the reviewers rather than toward numbers that look strict on paper. One principle is written into the code. A false rejection is the worst outcome, so borderline video goes to review instead of the bin, anomalous scores raise a warning rather than a rejection, and no automated check can reject a submission without clearing a confidence floor.

Captioning is where the length of these recordings becomes the whole problem. A caption has to reflect where an activity sits in the arc of a session rather than what happened to be in one sampled frame, which matters when somebody loads the washer, wanders off, and comes back twenty minutes later to unload it.

A caption has to say where an activity sits in a session, not what happened to be in one sampled frame

ONE UNSCRIPTED SESSION · 42 MIN

20 min between loading and unloading

loads the washer

elsewhere in the house

unloads the washer

folding

one sampled frame

0:00

42:00

FROM THE SAMPLED FRAME

“A person walks through a kitchen.”

True of that instant and useless as supervision — nothing in it says a wash cycle is running.

FROM THE WHOLE SESSION

“Loads the washer, leaves for twenty minutes, returns to unload it and folds the load.”

The activity placed in the arc of the session, with the order and the pauses intact.

For the wider picture of how footage like this gets prepared for world model training, our data platform team wrote up the full pipeline in World Model Data Pipeline.

What You Can Train On It

Two things that can make this corpus useful.

The first is language tied to real activity. A caption on a stock clip describes a scene. A caption over a continuous session describes what a person was doing and in what order, with the pauses, mistakes and corrections still in frame. That is the supervision instruction-conditioned models run on, and the interesting part is not the captioning. It is having ten thousand hours of unscripted first-person footage worth captioning in the first place.

The second is coverage of ordinary domestic environments. If you are training a household robot, a world model, or a video generation model that has to respect how objects behave when handled, the distance between your training data and a real kitchen matters. This is thousands of real kitchens at the hours of day when people are actually in them.

The raw tier is deliberately unopinionated, and it is there because our processing choices should not be forced on you. We clip, filter and caption to serve the projects we run, and any team with its own pipeline will want different boundaries, different thresholds and its own annotation schema. Releasing the footage as collected means you can start upstream of every decision we made.

Get The Data

The dataset is on Hugging Face under Apache 2.0, ungated.

Check out on Hugging Face

Following our June release of RekaCS2-10k, 10,000 hours of egocentric Counter-Strike 2 footage with per-frame action annotations and its accompanying open-source renderer, we are releasing RekaDaily-10k to advance progress in physical AI and foster an open research ecosystem.

Work With Us

Physical AI depends on footage captured in authentic, real-world environments. We operate a global network of over 100,000 collectors and are continuously scaling, making our data engine a powerful, extensible resource for your projects.

If this release aligns with your goals but you require a different setting, specific activities, a particular region, or specialized annotation, we can help. We provide custom data collection using the same proven pipeline. We have built bespoke datasets to spec for frontier research labs before, delivering completed projects with turnaround times measured in weeks.

This release is a snapshot of what our network produces daily. If you have specific requirements for data that doesn’t exist yet, whether for training or benchmarking, talk to us and we will build a programme to collect it.

Citation




Author