GPT-6 Astra, Claude Fable 5.1 and Fable 5 were each handed the same pair of robot arms and just two jobs: pick up the red block from the table and place it inside a bowl; grip a round puzzle piece by the knob at its center and insert it into the matching circular groove on the board. The independent evaluation shop Robocurve published the results on September 4, drawing an unusually clean capability line for the "physical hands" of frontier models, as a follow-up to their earlier Fable 5 versus 5.1 comparison.

19/20 vs 8/20: a lopsided block task

Each model ran 20 trials per task, 120 counted runs in total. On block-into-bowl, GPT-6 Astra completed 19 of 20, Claude Fable 5.1 managed 8, and the older Fable 5 just 1. Astra took about 2.5 minutes per trial to Fable 5.1's 6.8; estimated cost at list price was 0.94 USD per run against 2.12 USD, which Robocurve's chart labels as a 2.4x higher completion rate at 2.3x lower cost.

Scoring was not binary: a human grader scored each trial 0 to 4 by the highest stage reached, from "no purposeful approach" up to "placed in its final position", so failed runs still record how far they got. On the bowl task Astra's mean stage was 3.95, nearly always reaching the end; Fable 5.1 sat at 2.40 and Fable 5 at 1.30.

The output-token gap is more dramatic than the success rate

Astra emitted roughly 2,100 output tokens per run, against 12,900 for Fable 5.1 and 19,200 for Fable 5 — about one-sixth the output for more than twice the completions. The puzzle task shows the same pattern: 2,700 vs 10,500 tokens, at 1.36 USD vs 2.18 USD per run.

Precision insertion: all three models stall at the same step

The harder puzzle insertion is where the report gets interesting: Astra completed 2 of 20, Fable 5.1 also 2 of 20, and Fable 5 none at all. Astra reaches the groove and then stalls at the same final step Fable does. Grasping, moving and positioning all pass; what stalls is pressing the piece home. For a controller working only from three camera views and proprioceptive state, the missing ingredient is probably not more reasoning but something pure visual observation cannot supply.

How to read this evaluation

The limitations Robocurve itself lists deserve quoting: Astra's trials ran two days after the Fable trials, not interleaved; the bowl comparison was not run on the same rig; the grader knew which model was running, leaving room for unconscious bias; Anthropic requests went out without prompt caching while OpenAI auto-cached about a fifth of Astra's input without applying a discount — in the report's own words, Astra's cost is "if anything, overstated".

So keep the conclusion measured: on vision-guided grasping and coarse placement, Astra holds a simultaneous faster, cheaper and more accurate advantage; on fine insertion, frontier models still collectively fail. The value of this evaluation is not the ranking but that it pinpoints the next engineering problem to one specific step: the final insertion is not something a language model solves by thinking harder.

Reference: Robocurve, "GPT-6 Astra on robotic manipulation" — openai.robocurve.org/gpt-6-astra/