EGOCENTRIC INTELLIGENCE

EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos

Beyond seeing.Understand the tool.

EgoTools is a 100-hour dataset, a diagnostic benchmark, and an 8B reference model for reasoning about tools in real-world first-person video.

Cite this work BibTeX ↗

Shulin Tian1,2,*Junsu Kim1,3,*Shuai Liu1,*Hao Li1,*,★Yujiao Shen1Sihan Li1Zhe Yang1Yeongon Kim1Feiyu Li4Jialin Wu5Yichi Zhang1Wenhui Wang6Runmao Yao1Yuhao Dong1Zhaoxi Chen1,★Fangzhou Hong1,★Antonino Furnari7Jingkang Yang1Hongyuan Zhu2Ziwei Liu1,†,★

  1. 1S-Lab, Nanyang Technological University
  2. 2A*STAR
  3. 3KAIST
  4. 4PKU
  5. 5FDU
  6. 6School of Biological Sciences, Nanyang Technological University
  7. 7University of Catania

* Equal contribution† Corresponding author Ropedia Author

REAL CASE ROTATION / 01
GROUND → TRACK → INFER
KitchenSpatula

Transfer without piercing

  1. 01
    GroundLocate the tool
  2. 02
    TrackFollow its use
  3. 03
    InferExplain the choice
01 / DATA02 / BENCH03 / MODEL
FILM

THE FULL STORY

Watch tools become evidence.

The complete EgoTools project film—from real-world collection to tool-centric narration, grounding, evaluation, and learning.

01 / PROJECT FILMEgoTools in 100 hours of real-world tool use.
01

THE DATASET

Seven domains.
Real-world tool use.

From kitchens and labs to workshops and living spaces, EgoTools captures how people choose and use tools during real tasks.

Open 3D PlaygroundExperimental
First-person kitchen scene with a knife and cutting board
D-01Kitchen
First-person laboratory scene using a multichannel pipette
D-02Research Lab
First-person repair scene using pliers
D-03Repair Workshop
First-person classroom balloon rocket experiment
D-04Classroom
First-person craft scene sewing a plush object
D-05Craft
First-person office scene working at a desk
D-06Office
First-person household scene preparing coffee
D-07Household
Preview of the interactive Playground: a low-poly kitchen with selectable toolsInteractive demo

3D Playground

Experimental

Explore three rooms. Pick a tool.
Discover the video behind it.

Enter the playground
INTENT / ANNOTATIONWhy tools are chosen

EgoTools records the decision around a tool interaction—not just a list of low-level hand motions. Tool-centric narrations connect the selected tool and its properties to the task, the intended use, and the reason it was chosen over an alternative.

That structure preserves the task-state changes needed to study affordance, causality, and procedure across long, real-world episodes.

01 / INTENT

READ THE DECISION

Not just what happened.Why this tool, for this task, now.

A tool interaction becomes useful supervision when action is connected to purpose.

Narration interface and examples connecting tool choice to task intent
FIG / NARRATIONFrom visible action to tool choice and task intent.ANNOTATION / TOOL + PURPOSE + REASON
02

THE BENCHMARK

One thousand questions. No frame lookup.

Every question asks for evidence-grounded reasoning across tool choice, perception, procedure, or spatial relations.

40.34H RESERVED8-WAY MCQ4 TRACKS
AC
Affordance & Causality

Why this tool, and what changes because of it?

363
PG
Perception & Grounding

Which object, state, property, or evidence moment?

236
PD
Procedural Dynamics

How do steps, transitions, and outcomes connect?

222
SR
Spatial Reasoning

How do the hand, tool, and target align in 3D?

179

DEMO DIFFICULTY

Pick a level. Watch the clip. Choose an answer.

Tool identificationDEMO / Easy

What is the most accessible and convenient tool the person used to move the asparagus from the pan to the plate?

03

THE SIGNAL

Training on real tool use moves the needle.

The paper reports that EgoTools supervision lifts the 8B backbone from 50.0 to 60.9, a gain of 10.9 percentage points.

PAIRED COMPARISON / ACCURACY (%)Same backbone. Matched 64-frame protocol.
  1. Affordance & Causality (AC)
    +9.3
  2. Perception & Grounding (PG)
    +13.8
  3. Procedural Dynamics (PD)
    +11.0
  4. Spatial Reasoning (SR)
    −11.0

12.5% chance level · 8-way multiple choice

04 / LEADERBOARD

EgoTools-Bench Leaderboard

All results reported in the current arXiv manuscript, now in one sortable ranking. Choose any track to compare every model.

Have a new run? Send the checkpoint, input protocol, predictions, and reproduction notes.

Submit to EgoTools-Bench
Rank by
Accuracy (%)
RankModelCategoryEval. framesAudio
01Human ExpertHuman referenceFree-view✓82.185.283.382.783.2
02Gemini-3.1-ProProprietary1 fps✓69.751.772.174.966.9
03Gemini-3-FlashProprietary1 fps✓64.550.073.072.664.4
04EgoTools-8BOurs · Fine-tuned64—55.459.768.144.460.9
05Gemini-3.1-Flash-LiteProprietary1 fps✓58.144.559.958.755.4
06GLM4.1V-ThinkingOpen · Reasoning64—46.150.057.153.350.7
07Qwen3-VL-8B-InstructOpen · Instruct64—46.145.957.155.450.0
08InternVL3.5-8B-InstructOpen · Instruct64—45.845.953.453.348.7
09InternVL3-8B-InstructOpen · Instruct64—44.744.951.848.947.1
10Qwen2.5-Omni-7B-InstructOpen · Instruct64✓44.745.951.846.747.1
11Qwen3-VL-4B-InstructOpen · Instruct64—41.745.450.250.045.7
12Qwen3-VL-8B-ThinkingOpen · Reasoning512—43.641.848.255.445.7
13MiMo-VL-7B-RLOpen · Reasoning64—43.145.950.639.145.3
14Qwen2.5-VL-7B-InstructOpen · Instruct64—40.143.350.247.844.3
15MiMo-VL-7B-SFTOpen · Instruct64—41.444.949.440.244.2
16Qwen2-VL-7B-InstructOpen · Instruct64—46.635.646.247.844.2
17LLaVA-Video-7B-Qwen2Open · Instruct64—38.239.743.337.039.8
18LLaVA-OneVision-7BOpen · Instruct64—37.637.142.937.038.9
19Qwen3-VL-4B-ThinkingOpen · Reasoning512—37.636.636.441.337.4

Human Expert is shown as a free-viewing ceiling; model protocols remain visible through category, frame budget, and audio columns. Gemini models use 1 fps video.

A check means synchronized non-narration audio is provided. Narrations, dense captions, and annotation metadata are never used as evaluation inputs.

05

THE SYSTEM

From experience to intelligence.

A single traceable pipeline connects real-world collection to a model that can explain what it sees.

01 / Record

Real work.
First-person evidence.

646videos100.37 hours · 7 domains

People choose and use tools in everyday environments, with the task unfolding over time.

Input / egocentric video
02 / Annotate

Connect action, intent, and motion.

A / What happens361,332

Layered captions

Visible actions and task progress, from short moments to longer context.

5 sec1 min5 min
B / Why this tool6,519

Tool-centric narrations

The chosen tool, its intended use, and the reason it fits the task.

Tool + purpose + reason
C / Where it moves2D 3D

Grounded traces

Tool locations in the image, tracked motion, and reconstructed trajectories.

Space + time
03 / Separate source videosTwo resources. Disjoint recordings.

Benchmark source videos and their derived clips or annotations are excluded from instruction tuning.

Evaluate / benchmark poolEgoTools-Bench

Ask for evidence.

Reserved recordings support questions about tool choice, observed states, procedure, and spatial relationships.

1,000questions8 choices · 4 tracks
Affordance & causalityPerception & groundingProcedural dynamicsSpatial reasoning
Try the benchmark
Learn / training poolEgoTools-8B

Learn from tool use.

Training-pool videos and annotations form instruction examples for supervised video-language learning.

8Breference modelQwen3-VL-8B-Instruct backbone

The reference model is evaluated on the separate benchmark pool, under the same 64-frame protocol as its backbone.

See the measured gains
Visual evidenceInspect the annotation views

Representative views of grounding, temporal tracking, and 3D reconstruction.

EgoTools annotation view locating a kitchen tool in a first-person frame
GroundLocate the tool in the image.
EgoTools annotation view tracking a tool through kitchen footage
TrackFollow its position through time.
EgoTools annotation view showing a reconstructed tool trajectory in 3D
LiftInspect its motion in 3D space.
REFERENCE

CITATION

Cite EgoTools.

If you use EgoTools in your research, please cite our work.

BIBTEX / 2026

Copy the entry below or download a .bib file.

@misc{tian2026egotools,
  title = {{EgoTools}: Towards Tool-Centric Reasoning in Real-World Egocentric Videos},
  author = {Tian, Shulin and Kim, Junsu and Liu, Shuai and Li, Hao and Shen, Yujiao and
            Li, Sihan and Yang, Zhe and Kim, Yeongon and Li, Feiyu and Wu, Jialin and
            Zhang, Yichi and Wang, Wenhui and Yao, Runmao and Dong, Yuhao and Chen, Zhaoxi and
            Hong, Fangzhou and Furnari, Antonino and Yang, Jingkang and Zhu, Hongyuan and Liu, Ziwei},
  year = {2026},
  eprint = {2609.39378},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  doi = {10.48550/arXiv.2609.39378},
  url = {https://arxiv.org/abs/2609.39378}
}