MJAK egocentric hand-pose research with skeleton overlays
MJAK LabsPhysical AI Research

Research that reaches the robot

Physical AI. Measured. Ready to improve.

Building the evaluation and inference infrastructure for reliable robot intelligence.

MJAK LabsProblem

A working demo leaves the hard questions open

Can the result survive a new scene, a different arm, or a slower inference path?

80%

40 successful trials out of 50
A headline with substantial uncertainty

67.0%88.8%60%80%100%

Illustrative 95% Wilson interval. Trial count and test conditions change what a score can establish.

  • Evaluation uncertaintySmall trial counts can hide the difference between checkpoints.
  • Deployment mismatchWeights, normalization, cameras, and control rate must stay consistent.
  • Incomplete runtime evidenceTask success alone misses timing failures and intervention burden.
Research context: PhAIL, 2026. Interval is a worked example.
MJAK LabsWhy now

Open robot models make the systems gap visible

The research frontier now includes how policies are tested, served, and trusted.

2024

Accessible VLAs

OpenVLA releases an open 7B policy trained on 970k real robot demonstrations.

Researchers can adapt a foundation model instead of starting from scratch.

2025

Inference becomes control

Real-time chunking addresses pauses and discontinuities caused by inference delay.

Model execution and action execution need a shared timing contract.

2026

Reliability faces scrutiny

FailBench and PhAIL expose weaknesses in outcome judging and evaluation practice.

Deployment decisions need evidence beyond benchmark scores.

MJAK LabsResearch opportunities

Four open problems in physical AI

MJAK is studying failures in perception, policy evaluation, and remote inference.

GapUnanswered questionProposed research output
Reproducible VLA evaluationIs an apparent improvement larger than noise?Paired trials, confidence intervals, frozen policy contracts
VLM outcome judgingCan the judge see the evidence that proves success?Contact-aware evidence, calibrated abstention, human review
Perception validityHow much expected geometry is missing or wrong?Coverage-aware pose and depth benchmarks
Cloud inference timingCan the policy act before its observation becomes stale?Latency sweeps, deadline profiles, local fallback validation
Proposed MJAK agenda, informed by PhAIL, FailBench, and MJAK's published evaluations.
MJAK LabsSolution

One evidence trail across the physical AI workflow

MJAK is building toward a research cloud where every result retains its policy, conditions, and limits.

01 / DATA

Inspect

Perception coverage, geometry quality, and independent references.

02 / POLICY

Compare

Matched scenes, uncertainty, interventions, and reproducible contracts.

03 / INFERENCE

Qualify

Action freshness, tail latency, deadline misses, and fallback behavior.

04 / RELEASE

Verify

Scope-specific acceptance and evidence attached to the deployed artifact.

The research result should travel with the model.

Platform direction. The integrated workflow is proposed.
MJAK LabsWhat we have built

MJAK Evals, robotruth, and GR00T research.

Evaluation records

MJAK Evals

Multimodal evaluation of hand pose, segmentation, metric depth, mesh, and camera timing.

evals.mjak.in

Internal evaluator verification. Model quality gates remain open.

Open source / Apache 2.0

robotruth

Robot CI for learned policies, covering policy identity, statistics, episode evidence, drift, judging, and faults.

robotruth.mjak.in

Python library and CLI. Robot results are based on replayed logs.

Learned-policy simulation

GR00T research

Cloud inference and closed-loop control experiments using NVIDIA GR00T on an A40 GPU.

groot.mjak.in

V3 accepted for the declared GPU-local MuJoCo task.

Source: MJAK's public product and research pages, reviewed 3 October 2026.
MJAK LabsMJAK Evals

Evaluation exposes what the pipeline misses

A multimodal research pipeline with published conditions, negative results, and next steps.

Metric depth / mean per-frame AbsRel (%)

0102030Depth Pro16.197VDA streaming22.293VDA offline22.942

796 associations per model, one TUM scene, raw predictions. Lower is better. All three quality gates failed.

69/69

Focused evaluator tests
Separate software audit

32,179

Hand-pose baseline images
Coverage targets still open

RGB, registered sensor depth, and Depth Pro focal ablation review
Original exploratory focal-ablation review. Separate from the full baseline chart.

Run completion, prediction coverage, and physical accuracy are separate claims.

MJAK Evals, September 2026. Evaluator verification does not certify production readiness.
MJAK Labsrobotruth

Robot CI for learned policies

A Python library and CLI that turns checkpoints and rollout logs into evidence a lab can act on.

Contract

Verify weights, normalization, action semantics, cameras, and embodiment.

Statistics

Report intervals and determine whether checkpoint differences exceed noise.

Episodes

Record outcomes, interventions, failure classes, and provenance.

Fingerprint

Detect changes in camera geometry, lighting, and arm dynamics.

Judge

Fuse vision and motion evidence. Abstain when the evidence is insufficient.

Guard

Monitor execution faults such as stalls and erratic control.

Fits beside the existing training stack. Reads the artifacts it already produces.

robotruth v0.1.6 / Apache 2.0. Policy CI remains a simulation prototype.
MJAK Labsrobotruth evidence

Execution faults and task failures need different evidence

Ten action-stream detectors / real task failures

0.310.570.00.5 / chance1.0

A robot can move normally while the task goes wrong.

Observed across BotFails and DROID. This motivates research on visual and contact evidence.

361/362

Execution faults caught in recorded robot logs
Injected stalls and erratic control, within 0.1 to 0.2 s

0/338

Guard false alarms
95% interval: 0 to 1.1%

54

Public datasets ingested
Zero ingestion errors

Physical robot evidence is replay. No physical arm has been gated by the tool yet.

robotruth validation. Detection of injected faults does not establish task-failure detection.
MJAK LabsGR00T research

The higher score was not the accepted release

MJAK evaluated the policy against task, timing, and joint-limit gates together.

Held-out task success / 100 trials per candidate

0%50%100%V298% / Rejected: joint-limit gateV392% / Accepted: GPU-local MuJoCo scope

Whiskers: 95% Wilson intervals. V3 recorded zero joint-limit corrections. Eight genuine task failures remain.

Six original frames reconstructing a successful V3 learned-policy trial in MuJoCo
V3 success, seed 30238. Exact logged-command reconstruction of a learned-policy trial.

Real model inference on an A40 GPU. Robot, contacts, cameras, and world are simulated.

V2 / V3, 29 September 2026. Hardware safety and sim-to-real validation remain open.
MJAK LabsCloud inference gap

A task can succeed while the timing contract fails

MJAK's remote inference trials reveal the next systems research problem.

Observed p95 RPC latency / milliseconds

05001,000 msGPU-local208Internet baseline948Internet tuned809

Fallback ticks: 1.6% local, 20.1% Internet baseline, 16.4% tuned. Runs used different prefetch leads.

Cloud robotics needs
deadline-aware inference
and validated local fallback.

  • Observation ageMeasure how stale the input is when an action executes.
  • Tail latencyQualify p95 and p99 behavior alongside task outcome.
  • Control continuityKeep local execution predictable when cloud responses arrive late.
MJAK GR00T. WAN: two 5-trial runs, simulated robot, real Internet route. Task-functional, not timing-qualified.
MJAK LabsCloud and inference direction

A research cloud with an explicit control boundary

The cloud supplies compute and evidence. The robot retains local control and stop behavior.

Research cloud

Experiments & evaluation

Versioned data and policies
Reproducible GPU runs
Paired scene batteries
Budgets and audit records

Inference service

Policy serving

Model-specific observation adapters
Asynchronous action chunks
Deadline and freshness profiles
Latency and resource telemetry

Robot or simulator

Local execution

Action queue and limits
Validated fallback behavior
Calibration and hardware interlocks
Episode evidence capture

A shared record binds the checkpoint, normalizer, camera setup, controller profile, and acceptance report.

Research objective: lower cost per successful task without losing reliability or control continuity.

Proposed MJAK platform. Informed by LeRobot asynchronous inference and RTC.
MJAK LabsVLM research direction

The judge needs evidence of the physical outcome

0.77

Best mean balanced accuracy
Across FailBench

≤0.60

Best balanced accuracy
Contact-rich assembly

External benchmark findings, not MJAK results. Different scopes. Thirteen detectors tested.

  • Outcome-specific evidenceUse the frames and views that can prove the task completed.
  • Calibrated abstentionReport answer coverage with error, then route ambiguity to humans.
  • Physical contextInvestigate temporal, depth, contact, and intervention signals.
Original MJAK known-pose mesh diagnostic showing RGB, mesh, raycast depth, and sensor depth
MJAK's known-pose geometry diagnostic. Ground-truth poses supplied. Metric reconstruction remains unaccepted.

Proposed research: determine when a visual judgment is supported by physical evidence.

FailBench, 2026 / MJAK Evals. Research agenda, not a solved outcome judge.
MJAK LabsMarket context

The robot base is growing. So is the need to validate intelligence.

Global industrial robot installations / thousands

8004000542>6006558062024202520262029ForecastForecast

2025 bar uses the reported 600k lower bound. Outline bars show IFR forecasts.

5M

Industrial robots operating globally in 2025

~10.5k

India installations in 2025
15% annual growth

+9%

Global operational stock
Annual growth in 2025

Hardware adoption is a demand signal. It is not a revenue estimate for MJAK's tooling.

MJAK LabsTarget users

The first users already have a policy and a reliability question

Start with research labs and robotics teams that need repeated, comparable experiments.

01

Embodied AI labs

Compare VLAs and VLMs with shared protocols, retained artifacts, and meaningful uncertainty.

Initial use: reproduce a result and diagnose why it changed.

02

Robotics startups

Check policy releases and profile inference without assembling an entire platform team.

Initial use: qualify a checkpoint for one robot and one task.

03

Automation teams

Evaluate changes in camera setup, calibration, intervention burden, and operating conditions.

Expansion use: validate a change before a broader rollout.

MJAK LabsEcosystem position

Strong building blocks. Room for an evidence layer.

Existing layerEstablished contributionMJAK's proposed focus
NVIDIA OSMOPhysical AI workflow and compute orchestrationConnect run artifacts to policy and inference acceptance
LeRobotOpen robot learning stack and asynchronous inferenceVerify deployment configuration and timing under shift
PhAILReal-robot evaluation with distributional methodologyBring rigorous comparisons into recurring policy CI
FailBenchCross-source evaluation of robot outcome judgesResearch calibrated judgments with physical evidence
MJAK LabsPublished evals, robotruth, and GR00T experimentsAn auditable record across perception, policy, and inference

Build on open models and runtimes. Make their operating limits measurable.

MJAK LabsResearch method

Each failure should produce a better experiment

Reusable protocols let another lab test the same claim.

01 / DEFINE

Freeze the claim

Declare the task, embodiment, model, controller, and pass conditions.

02 / MEASURE

Stress the setup

Vary scenes, calibration, timing, and policies with retained negative evidence.

03 / EXPLAIN

Locate the failure

Separate perception, task execution, configuration, and infrastructure causes.

04 / REPRODUCE

Publish the record

Release protocols and evidence for outsiders to reproduce and challenge.

Protocols

Clear comparison rules

Failure corpora

Known operating limits

Adapters

Repeatable model integration

Evidence

Results others can inspect

Proposed research process. Published negative evidence: robotruth and GR00T.
MJAK LabsAdoption and sustainability

Open tools now. Managed experiments next.

Open

Tools & protocols

Use robotruth locally, inspect evaluation records, and reproduce experiments.

Success measure: another lab reproduces a result.

Cloud

Managed experiments

Proposed GPU-backed evaluation runs, policy endpoints, private artifacts, and team access.

Potential revenue: compute usage and managed service fees.

Partner

Joint research

Collaborate with labs and robot teams on new embodiments, failure data, and inference studies.

Potential support: sponsored research and shared infrastructure.

MJAK LabsRoadmap

What must be proven next

StageWorkProof required to advance
01 / ReproduceIndependent installations, second policy, broader task batteriesExternal reproductions and measured null comparisons
02 / GroundPhysical in-loop tests and independent perception referencesScoped hardware evidence, calibration, intervention logs
03 / Qualify inferenceNetwork jitter, observation staleness, chunking, fallback sweepsTiming-qualified operating profiles and task reliability
04 / Offer the cloudManaged experiments, team access, retained release evidenceRepeat partner use and measured cost per successful task
MJAK LabsFounders

Meet the team

Akhilesh Chandra

Akhilesh Chandra

Founder

Machine learning, perception pipelines, evaluation systems, and cloud engineering.

  • ML Engineer, Deccan AIBuilt a synthetic database and reinforcement learning gyms
  • First Engineer, Peak AI RoboticsBuilt annotation and labelling pipelines
  • Computer Vision Engineer, Asai LabsDesigned the architecture for a shelf fill rate detection model
  • Activate VC AI Fellow
Shreyash Raut

Shreyash Raut

Co-founder / Robotics Engineer

Robotics engineering and experience across early-stage teams.

  • Founding Robotics Engineer, MyTron Labs
  • Operations Intern, Human Archive (YC W26)
  • Founding Member, DatraAI
  • Product Manager Intern, stealth startup
MJAK LabsSources and scope

The sources behind this deck

Public research records and primary sources, reviewed 3 October 2026.

  1. MJAK EvalsPerception metrics, evaluator verification, original diagnostic imagery
  2. robotruthModules, validation cohorts, correction, and current limitations
  3. MJAK GR00T researchComparison and V1, V2, V3 records. Learned inference in MuJoCo
  4. Akhilesh ChandraProfessional profile. Shreyash biography supplied by the team
  5. IFR World Robotics 20262025 adoption and forecasts. 2024 figure from the 2025 report
  1. OpenVLA, 2024Open foundation models for robot policy adaptation
  2. Real-Time Execution of Action Chunking Flow PoliciesInference latency and asynchronous action execution
  3. PhAIL, 2026Real-robot evaluation and distributional methodology
  4. FailBench, 2026Cross-source VLM outcome judging and contact evidence gaps
  5. NVIDIA OSMO / LeRobot async inferenceExisting orchestration and serving building blocks
Cloud architecture, target segments, sustainability model, and roadmap describe proposed work.
MJAK LabsThank you

Research partners. Robot teams. Infrastructure builders.

Let's make physical AI easier to trust.

Bring a working policy, a robot platform, or a difficult failure case.
Help us build the evidence and inference systems around it.

MJAK Labs | Physical AI Research