Research that reaches the robot
Physical AI. Measured. Ready to improve.
Building the evaluation and inference infrastructure for reliable robot intelligence.
A working demo leaves the hard questions open
Can the result survive a new scene, a different arm, or a slower inference path?
40 successful trials out of 50
A headline with substantial uncertainty
Illustrative 95% Wilson interval. Trial count and test conditions change what a score can establish.
- Evaluation uncertaintySmall trial counts can hide the difference between checkpoints.
- Deployment mismatchWeights, normalization, cameras, and control rate must stay consistent.
- Incomplete runtime evidenceTask success alone misses timing failures and intervention burden.
Open robot models make the systems gap visible
The research frontier now includes how policies are tested, served, and trusted.
Accessible VLAs
OpenVLA releases an open 7B policy trained on 970k real robot demonstrations.
Researchers can adapt a foundation model instead of starting from scratch.
Inference becomes control
Real-time chunking addresses pauses and discontinuities caused by inference delay.
Model execution and action execution need a shared timing contract.
Reliability faces scrutiny
FailBench and PhAIL expose weaknesses in outcome judging and evaluation practice.
Deployment decisions need evidence beyond benchmark scores.
Four open problems in physical AI
MJAK is studying failures in perception, policy evaluation, and remote inference.
| Gap | Unanswered question | Proposed research output |
|---|---|---|
| Reproducible VLA evaluation | Is an apparent improvement larger than noise? | Paired trials, confidence intervals, frozen policy contracts |
| VLM outcome judging | Can the judge see the evidence that proves success? | Contact-aware evidence, calibrated abstention, human review |
| Perception validity | How much expected geometry is missing or wrong? | Coverage-aware pose and depth benchmarks |
| Cloud inference timing | Can the policy act before its observation becomes stale? | Latency sweeps, deadline profiles, local fallback validation |
One evidence trail across the physical AI workflow
MJAK is building toward a research cloud where every result retains its policy, conditions, and limits.
01 / DATA
Inspect
Perception coverage, geometry quality, and independent references.
02 / POLICY
Compare
Matched scenes, uncertainty, interventions, and reproducible contracts.
03 / INFERENCE
Qualify
Action freshness, tail latency, deadline misses, and fallback behavior.
04 / RELEASE
Verify
Scope-specific acceptance and evidence attached to the deployed artifact.
The research result should travel with the model.
MJAK Evals, robotruth, and GR00T research.
Evaluation records
MJAK Evals
Multimodal evaluation of hand pose, segmentation, metric depth, mesh, and camera timing.
evals.mjak.inInternal evaluator verification. Model quality gates remain open.
Open source / Apache 2.0
robotruth
Robot CI for learned policies, covering policy identity, statistics, episode evidence, drift, judging, and faults.
robotruth.mjak.inPython library and CLI. Robot results are based on replayed logs.
Learned-policy simulation
GR00T research
Cloud inference and closed-loop control experiments using NVIDIA GR00T on an A40 GPU.
groot.mjak.inV3 accepted for the declared GPU-local MuJoCo task.
Evaluation exposes what the pipeline misses
A multimodal research pipeline with published conditions, negative results, and next steps.
Metric depth / mean per-frame AbsRel (%)
796 associations per model, one TUM scene, raw predictions. Lower is better. All three quality gates failed.
Focused evaluator tests
Separate software audit
Hand-pose baseline images
Coverage targets still open

Run completion, prediction coverage, and physical accuracy are separate claims.
Robot CI for learned policies
A Python library and CLI that turns checkpoints and rollout logs into evidence a lab can act on.
Contract
Verify weights, normalization, action semantics, cameras, and embodiment.
Statistics
Report intervals and determine whether checkpoint differences exceed noise.
Episodes
Record outcomes, interventions, failure classes, and provenance.
Fingerprint
Detect changes in camera geometry, lighting, and arm dynamics.
Judge
Fuse vision and motion evidence. Abstain when the evidence is insufficient.
Guard
Monitor execution faults such as stalls and erratic control.
Fits beside the existing training stack. Reads the artifacts it already produces.
Execution faults and task failures need different evidence
Ten action-stream detectors / real task failures
A robot can move normally while the task goes wrong.
Observed across BotFails and DROID. This motivates research on visual and contact evidence.
Execution faults caught in recorded robot logs
Injected stalls and erratic control, within 0.1 to 0.2 s
Guard false alarms
95% interval: 0 to 1.1%
Public datasets ingested
Zero ingestion errors
Physical robot evidence is replay. No physical arm has been gated by the tool yet.
The higher score was not the accepted release
MJAK evaluated the policy against task, timing, and joint-limit gates together.
Held-out task success / 100 trials per candidate
Whiskers: 95% Wilson intervals. V3 recorded zero joint-limit corrections. Eight genuine task failures remain.

Real model inference on an A40 GPU. Robot, contacts, cameras, and world are simulated.
A task can succeed while the timing contract fails
MJAK's remote inference trials reveal the next systems research problem.
Observed p95 RPC latency / milliseconds
Fallback ticks: 1.6% local, 20.1% Internet baseline, 16.4% tuned. Runs used different prefetch leads.
Cloud robotics needs
deadline-aware inference
and validated local fallback.
- Observation ageMeasure how stale the input is when an action executes.
- Tail latencyQualify p95 and p99 behavior alongside task outcome.
- Control continuityKeep local execution predictable when cloud responses arrive late.
A research cloud with an explicit control boundary
The cloud supplies compute and evidence. The robot retains local control and stop behavior.
Research cloud
Experiments & evaluation
Versioned data and policies
Reproducible GPU runs
Paired scene batteries
Budgets and audit records
Inference service
Policy serving
Model-specific observation adapters
Asynchronous action chunks
Deadline and freshness profiles
Latency and resource telemetry
Robot or simulator
Local execution
Action queue and limits
Validated fallback behavior
Calibration and hardware interlocks
Episode evidence capture
Research objective: lower cost per successful task without losing reliability or control continuity.
The judge needs evidence of the physical outcome
Best mean balanced accuracy
Across FailBench
Best balanced accuracy
Contact-rich assembly
External benchmark findings, not MJAK results. Different scopes. Thirteen detectors tested.
- Outcome-specific evidenceUse the frames and views that can prove the task completed.
- Calibrated abstentionReport answer coverage with error, then route ambiguity to humans.
- Physical contextInvestigate temporal, depth, contact, and intervention signals.

Proposed research: determine when a visual judgment is supported by physical evidence.
The robot base is growing. So is the need to validate intelligence.
Global industrial robot installations / thousands
2025 bar uses the reported 600k lower bound. Outline bars show IFR forecasts.
Industrial robots operating globally in 2025
India installations in 2025
15% annual growth
Global operational stock
Annual growth in 2025
Hardware adoption is a demand signal. It is not a revenue estimate for MJAK's tooling.
The first users already have a policy and a reliability question
Start with research labs and robotics teams that need repeated, comparable experiments.
Embodied AI labs
Compare VLAs and VLMs with shared protocols, retained artifacts, and meaningful uncertainty.
Initial use: reproduce a result and diagnose why it changed.
Robotics startups
Check policy releases and profile inference without assembling an entire platform team.
Initial use: qualify a checkpoint for one robot and one task.
Automation teams
Evaluate changes in camera setup, calibration, intervention burden, and operating conditions.
Expansion use: validate a change before a broader rollout.
Strong building blocks. Room for an evidence layer.
| Existing layer | Established contribution | MJAK's proposed focus |
|---|---|---|
| NVIDIA OSMO | Physical AI workflow and compute orchestration | Connect run artifacts to policy and inference acceptance |
| LeRobot | Open robot learning stack and asynchronous inference | Verify deployment configuration and timing under shift |
| PhAIL | Real-robot evaluation with distributional methodology | Bring rigorous comparisons into recurring policy CI |
| FailBench | Cross-source evaluation of robot outcome judges | Research calibrated judgments with physical evidence |
| MJAK Labs | Published evals, robotruth, and GR00T experiments | An auditable record across perception, policy, and inference |
Build on open models and runtimes. Make their operating limits measurable.
Each failure should produce a better experiment
Reusable protocols let another lab test the same claim.
01 / DEFINE
Freeze the claim
Declare the task, embodiment, model, controller, and pass conditions.
02 / MEASURE
Stress the setup
Vary scenes, calibration, timing, and policies with retained negative evidence.
03 / EXPLAIN
Locate the failure
Separate perception, task execution, configuration, and infrastructure causes.
04 / REPRODUCE
Publish the record
Release protocols and evidence for outsiders to reproduce and challenge.
Clear comparison rules
Known operating limits
Repeatable model integration
Results others can inspect
Open tools now. Managed experiments next.
Tools & protocols
Use robotruth locally, inspect evaluation records, and reproduce experiments.
Success measure: another lab reproduces a result.
Managed experiments
Proposed GPU-backed evaluation runs, policy endpoints, private artifacts, and team access.
Potential revenue: compute usage and managed service fees.
Joint research
Collaborate with labs and robot teams on new embodiments, failure data, and inference studies.
Potential support: sponsored research and shared infrastructure.
What must be proven next
| Stage | Work | Proof required to advance |
|---|---|---|
| 01 / Reproduce | Independent installations, second policy, broader task batteries | External reproductions and measured null comparisons |
| 02 / Ground | Physical in-loop tests and independent perception references | Scoped hardware evidence, calibration, intervention logs |
| 03 / Qualify inference | Network jitter, observation staleness, chunking, fallback sweeps | Timing-qualified operating profiles and task reliability |
| 04 / Offer the cloud | Managed experiments, team access, retained release evidence | Repeat partner use and measured cost per successful task |
Meet the team

Akhilesh Chandra
Founder
Machine learning, perception pipelines, evaluation systems, and cloud engineering.
- ML Engineer, Deccan AIBuilt a synthetic database and reinforcement learning gyms
- First Engineer, Peak AI RoboticsBuilt annotation and labelling pipelines
- Computer Vision Engineer, Asai LabsDesigned the architecture for a shelf fill rate detection model
- Activate VC AI Fellow
SRShreyash Raut
Co-founder / Robotics Engineer
Robotics engineering and experience across early-stage teams.
- Founding Robotics Engineer, MyTron Labs
- Operations Intern, Human Archive (YC W26)
- Founding Member, DatraAI
- Product Manager Intern, stealth startup
The sources behind this deck
Public research records and primary sources, reviewed 3 October 2026.
- MJAK EvalsPerception metrics, evaluator verification, original diagnostic imagery
- robotruthModules, validation cohorts, correction, and current limitations
- MJAK GR00T researchComparison and V1, V2, V3 records. Learned inference in MuJoCo
- Akhilesh ChandraProfessional profile. Shreyash biography supplied by the team
- IFR World Robotics 20262025 adoption and forecasts. 2024 figure from the 2025 report
- OpenVLA, 2024Open foundation models for robot policy adaptation
- Real-Time Execution of Action Chunking Flow PoliciesInference latency and asynchronous action execution
- PhAIL, 2026Real-robot evaluation and distributional methodology
- FailBench, 2026Cross-source VLM outcome judging and contact evidence gaps
- NVIDIA OSMO / LeRobot async inferenceExisting orchestration and serving building blocks
Research partners. Robot teams. Infrastructure builders.
Let's make physical AI easier to trust.
Bring a working policy, a robot platform, or a difficult failure case.
Help us build the evidence and inference systems around it.