The flight simulatorfor home-health AI.

Robots, smart glasses and care apps are moving into patients' homes. Today the only test is a real patient.

HealthDojo turns clinical guidelines into synthetic homes with exact labels, so your team finds the edge cases and trains on them before a patient ever does.

Cluttered staircase · stairs-base0-STAIR-03
Live 3D · Gaussian splat · walk it yourself →
scroll
Why now
Healthcare AI is moving into the home. We built it a clinically grounded place to train.
01 · e.g. 1X NEO, Figure

Humanoid & home robots

About to share homes with older adults. They have to know a towel bar is not a grab bar.

02 · e.g. Meta Ray-Ban Display

Smart glasses & wearables

Seeing the home through the user's eyes, in every room the patient walks through.

03 · discharge, monitoring, caregiving

Home-health & care apps

A wave of apps assessing homes from a phone camera, all judging safety from pixels.

Every one needs to be tested before it meets a patient. Today the only test is real patients. HealthDojo is the dojo: validated, guideline-grounded environments where models and apps train and get pressure-tested.
How it works · one real row, end to end
A guideline line becomes a rubric row # becomes a labelled home ▣ becomes a verdict ✗
01Guideline
CDC

CDC STEADI · Check for Safety

U.S. Centers for Disease Control and Prevention (CDC)
“Stairs and steps: keep objects off the stairs; fix loose or uneven steps; make sure carpet is firmly attached to every step, or remove it and put non-slip rubber treads on the stairs; handrails on both sides, as long as the stairs; fix loose handrails.”
https://www.cdc.gov/steadi/pdf/STEADI-Brochure-CheckForSafety-508.pdf
Reference document. HealthDojo is not affiliated with or endorsed by CDC.
02Rubric

STAIR-02 Handrail loose, short, or not graspable

type
adequacy/measurement
room
stairs
severity
high Loose/broken or ungraspable rail on only rail
ICD-10
W10.8
Fall on and from other stairs and steps
cites
CDC: Stairs & Steps, HSSAT: Stairs, WeHSA: Steps/Stairs
03Synthetic home
TRUTH · STAIR-02
label written before the pixels · judge-verified ✓
Left handrail shortened to a small stub near the bottom, so it no longer runs the length of the stair. The rest of the scene is unchanged.
04Verdict
HIT
Claude Opus 5.5
flagged STAIR-02, STAIR-01, LIV-01
MISS
Grok 4.6
flagged STAIR-01, STAIR-05, LIV-01, not STAIR-02
MISS
Nova Pro
flagged STAIR-01, not STAIR-02
8 of 14 models caught STAIR-02
Guidelines in. Benchmarks out.Paste any home-safety guideline and watch it compile into a rubric, scenes and a scored benchmark.Watch models walk it →See the compiled guidelines →
The curriculum

Five levels. Each one harder.

Every level ships as labelled images and walkable 3D worlds.

Every image is built from the rubric: label first, then paint, then verify.

L1

Obvious

One clear hazard, clean room, good light.

Hazards
Look-alikesnone
Conditionsgood light
Image3D →
L2

Subtle

Judge adequacy: a towel bar, a short rail, a low-contrast mat.

Hazards
Look-alikesnone
Conditionsgood light
Image3D →
L3

Compound

Two hazards plus a safe look-alike that must not be flagged.

Hazards
Look-alikes
Conditionsgood light
Image3D →
L4

Cluttered

Three or four hazards in dim, noisy, unfamiliar rooms.

Hazards3-4
Look-alikes
Conditionsdim, night, soft focus
Image3D →
L5

Adversarial

Safe-but-scary rooms next to dense hazard rooms, in poor light.

Hazards0 or 4+
Look-alikes2-3
Conditionspoor light
Image3D →
Images and 3D worlds

One label, two formats.

Every scene ships as a labelled image and a walkable 3D world.
Labelled imageWalkable 3D world
Simulators

Every guideline is a new simulator.

Open the library →

NextPressure injuriesNPIAP
NextMedication safetyAHRQ / ISMP
NextPediatric home injuryAAP
NextSmoke & CO alarmsNFPA 72
HomeBench · open benchmark

Leaderboard

14 models on 43 verified scenes. Score blends recall and false alarms.
#ModelProviderScoreRecallFalse alarms
1Claude Opus 5.5Anthropic
0.92
100%17%
2GPT-5.6 SolOpenAI
0.90
97%18%
3GPT-5.6 TerraOpenAI
0.84
83%14%
4Kimi K3Moonshot
0.84
93%24%
5Grok 4.6xAI
0.81
73%11%
6Claude Sonnet 5Anthropic
0.81
87%26%
7Llama 4 MaverickMeta
0.79
83%26%
8Nova ProAmazon
0.77
77%22%
9Qwen3-VLAlibaba
0.75
84%34%
10Nova 2 LiteAmazon
0.73
67%20%
11Gemma 3 27BGoogle
0.70
81%40%
12Claude Haiku 4.5Anthropic
0.70
83%43%
13Nemotron Nano VLNVIDIA
0.70
90%50%
14Mistral Large 3Mistral
0.62
84%59%
Who it's for

For teams putting AI in the home.Find the edge cases. Train on them.

Home robots

Humanoids that share the hallway

Train hazard perception in walkable 3D homes before meeting a walker.

Smart glasses & wearables

Vision on the patient's face

Score what the device notices, and what it misses, from the user's view.

Home-health & care apps

Edge cases before the pilot

Find blind spots on synthetic homes instead of real patients.

What you get
A private simulatorCompiled from your guideline, your rooms, your hazard list.
Hard-example training packsEvery miss, turned into labelled data at the level where it breaks.
Re-run on every releaseNew checkpoint, same simulator: regressions surface first.
Every failure we find becomes new training data, and the simulator gets harder.
Later: hospitals and payers under CJR-X, the rating they require.
The stack
Built in one day on AWS.
Built in one day at the Healthcare AI Hackathon, AWS Builder Loft, San Francisco. HomeBench is open on GitHub.
Powered byAWSAWSAnthropicAnthropicGoogle GeminiGoogle GeminiStability AIStability AIWorld LabsWorld LabsElevenLabsElevenLabsOpenAIOpenAIGitHubGitHub
AWS
Model access
AWS Bedrock

13 of 14 models we tested run on Bedrock: Claude, OpenAIOpenAI GPT-5.6 via Bedrock, Nova, Llama 4, Qwen3-VL, Mistral, Gemma, Kimi, Grok, Nemotron.

ClaudeAnthropic
Rubric + judge
Claude Sonnet 5 and Claude Opus 5.5

Sonnet 5 on Bedrock drafts cited rubric rows; the verifier checks every scene before it counts. Grading against labels is deterministic, no LLM grades the answers.

Stability AIGoogle Gemini
Image generation
Stability AI on Bedrock + Gemini 3 Pro Image

Clean rooms, then hazards painted in one label at a time.

World Labs
3D worlds
World Labs Marble

Gaussian-splat homes with collision meshes, walkable in the browser.

Agent environment
Custom look-around harness

Models choose where to look, then flag what they find.

ElevenLabs
Voice
ElevenLabs

Narrates each walkthrough step.

AWSGitHub
Hosting
Amazon S3 + GitHub

Static site on S3; code and benchmark open on GitHub.

14models tested
2guidelines compiled
134labelled homes
11walkable 3D worlds
Synthetic data only. Rubrics pending OT/PT review. Sim-to-real validation is next.
Looking for design partners: fly your model through the simulator.
Guideline libraryGitHub ↗