Maccark
Talk to us

Real-world evaluations of physical AI

We measure and report frontier robotics capabilities to the public.

Talk to usWhy it matters

Opus 5 on block stacking


Stack the red block on top of the blue block.

LLM inference speed over time


LLM inference speed over time

Test-time scaling in robotics capabilities


Test-time scaling in robotics capabilities

Opus 5 and the Tower of Hanoi


Lift the finished slice out of the toaster…

Opus 5 on block stacking


Stack the red block on top of the blue block.

LLM inference speed over time


LLM inference speed over time

Test-time scaling in robotics capabilities


Test-time scaling in robotics capabilities

Opus 5 and the Tower of Hanoi


Lift the finished slice out of the toaster…
Backed by NVIDIA InceptionAnthropic for Open SourceFeatured by Google for Developers

Opus 5 and the Tower of Hanoi

Complete the Tower of Hanoi puzzle by moving all disks to the rightmost peg.

With support from experts at:

MIT
STANFORD
HARVARD
PRINCETON
CALTECH
UW
METR
MATS

Independent analysis of robotics capabilities

Test-time scaling in robotics capabilities

Robotics capabilities of LLMs scale with thinking effort.

Test-time scaling in robotics capabilities
Source: Maccark, Opus 5 report (Aug 2026)Methodology
Why it matters

A technology this big must be measured in the open

02 h4 h6 h8 h10 h12 h14 h16 h18 h20192020202120222023202420252026Task length at 50% success (human hours)Model release dateGPT-2 (2019-02-14): 0.1 minGPT-3 (davinci-002) (2020-05-28): 0.1 minGPT-3.5 Turbo Instruct (2022-03-15): 0.6 minGPT-4 (0314) (2023-03-14): 4.0 minGPT-4 (1106) (2023-11-06): 4.0 minClaude 3 Opus (2024-03-04): 4.0 minGPT-4 Turbo (2024-04-09): 3.7 minGPT-4o (2024-05-13): 7.0 minClaude 3.5 Sonnet (Jun 2024) (2024-06-20): 11.4 mino1-preview (2024-09-12): 20.3 minClaude 3.5 Sonnet (Oct 2024) (2024-10-22): 20.5 mino1 (2024-12-05): 38.8 minClaude 3.7 Sonnet (2025-02-24): 1.0 ho3 (2025-04-16): 2.0 hClaude Opus 4 (2025-05-22): 1.7 hClaude Opus 4.1 (2025-08-05): 1.7 hGPT-5 (2025-08-07): 3.4 hGemini 3 Pro (2025-11-18): 3.7 hGPT-5.1-Codex-Max (2025-11-19): 3.7 hClaude Opus 4.5 (2025-11-24): 4.9 hGPT-5.2 (2025-12-11): 5.9 hClaude Opus 4.6 (2026-02-05): 12.0 hGPT-5.3-Codex (2026-02-05): 5.8 hGemini 3.1 Pro (2026-02-19): 6.4 hGPT-5.4 (2026-03-05): 5.7 hClaude Mythos Preview (early) (2026-04-07): 17.4 h

Source: METR, Task-Completion Time Horizons of Frontier AI Models

The stakes

Robotics capabilities are progressing rapidly

Frontier labs are racing to develop general-purpose robots in the next two years.

The problem

No one knows where the frontier is

The Internet is filled with cherry-picked demo videos. No continuous, standardized evaluations for robotics exist.

Our answer

We are an independent evaluator for robotics

We measure robotics capabilities in the real world, independent of any agenda.

Our work

Open tools for open evaluation

Open source

Our tools and datasets are public so the community can verify and build on our work.

Contribute
Inspect Robots UI
Product

Inspect Robots

A developer toolkit for evaluating robot behavior, gathering demonstration data, and benchmarking physical AI.

  • Live robot session recording
  • Multi-model task benchmarking
  • Shareable eval reports
About us

Maccark is a Public Benefit Corporation helping society understand the frontier of physical AI.

Our mission

Maccark is a Public Benefit Corporation bound by law to serve the public good. General-purpose robots may arrive within years, with profound implications for the economy and the labor market.

We build open-source tools and independent benchmarks to measure robotics capabilities. We report progress to the public rigorously and neutrally.

What experts say

Words from the experts

Researchers, engineers, and forecasters on Maccark.

"Supply chain automation timelines are a crucial input to ASI timelines in hardware-dependent AI takeoff scenarios, as well as forecasting progress in AI military technologies. This kind of benchmarking effort could help ground those forecasts."

Ryan Kidd
Ryan KiddCEO & Co-Founder at MATS Research

"Timelines to robot automation is an important input into our models of takeoff, yet it is one that I (and in my experience, many other people) have a lot of uncertainty about. This project appears well-scoped to make progress on this."

Gabe Wu
Gabe WuAlignment Researcher at OpenAI

"There is a need for good robotics benchmarks to capture capability improvements that may emerge in the coming years. This project could mark a strong contribution to this space."

Julian Jacobs
Julian JacobsResearch Scientist and Economist at Google DeepMind

"Supply chain automation timelines are a crucial input to ASI timelines in hardware-dependent AI takeoff scenarios, as well as forecasting progress in AI military technologies. This kind of benchmarking effort could help ground those forecasts."

Ryan Kidd
Ryan KiddCEO & Co-Founder at MATS Research

"Timelines to robot automation is an important input into our models of takeoff, yet it is one that I (and in my experience, many other people) have a lot of uncertainty about. This project appears well-scoped to make progress on this."

Gabe Wu
Gabe WuAlignment Researcher at OpenAI

"There is a need for good robotics benchmarks to capture capability improvements that may emerge in the coming years. This project could mark a strong contribution to this space."

Julian Jacobs
Julian JacobsResearch Scientist and Economist at Google DeepMind

"We've all seen the demo where a robot does a backflip or jumps rope, and then you try to use it for anything real and, in the best case, it doesn't break your own robot. That gap is exactly why independent measurement matters."

Liane Galanti
Liane GalantiPhD Student in Computer Science at Princeton

"Measuring robotics capabilities over time seems like a very important input for forecasting AI takeoff speeds. Currently, there are basically no widely-used high-quality robotics benchmarks, and additional high-quality evals here seem very valuable."

Nikola Jurkovic
Nikola JurkovicMember of Technical Staff at METR

"Physical AI may be one of the most transformative technologies in human history, and yet we cannot say with any confidence what robots can do today, how fast it is changing, or when it will start to affect human labour markets."

Sebastian Sartor
Sebastian SartorPhD Student in Mechanical Engineering at MIT

"We've all seen the demo where a robot does a backflip or jumps rope, and then you try to use it for anything real and, in the best case, it doesn't break your own robot. That gap is exactly why independent measurement matters."

Liane Galanti
Liane GalantiPhD Student in Computer Science at Princeton

"Measuring robotics capabilities over time seems like a very important input for forecasting AI takeoff speeds. Currently, there are basically no widely-used high-quality robotics benchmarks, and additional high-quality evals here seem very valuable."

Nikola Jurkovic
Nikola JurkovicMember of Technical Staff at METR

"Physical AI may be one of the most transformative technologies in human history, and yet we cannot say with any confidence what robots can do today, how fast it is changing, or when it will start to affect human labour markets."

Sebastian Sartor
Sebastian SartorPhD Student in Mechanical Engineering at MIT
Contact

We work with labs of all sizes

You're interested in: