PointArena 2026

Efficient Visual Pointing for Embodied AI

Agent-Driven Data Synthesis · AttnRes Fine-Tuning · Iterative Correction · Molmo2-8B

2nd Place — 77.2% Overall
93.9%
Affordance
82.6%
Spatial
78.2%
Reasoning
63.0%
Steerability
Motivation

Three Bottlenecks

1. Data Scarcity

Steerable has ZERO public training data.

→ 55K processed, 37.6K trainable

2. Architecture Gap

Standard LoRA weakly couples coords with semantics.

→ AttnRes: cross-block attention

3. One-Shot Prediction

Single forward pass, no recovery.

→ Correction-based training

Molmo2-8B Zero-Shot

ModelAfford.SpatialReason.CountingSteer.Overall
Molmo2-8B85.9%76.9%77.2%73.0%50.5%72.7%
Innovation 1

Agent-Driven Data Synthesis Pipeline

22.3M
Raw QA
~2.5M
Has Coords
55,372
Processed
37,574
Trainable
3,815
Early Subset
0
Human
22.3M Raw QA Filter → ~2.5M by_source/*.jsonl Pipeline A: Gemini API (5-Class)Gatekeeper→5-way Classify→MP→Rewritegemini-3-flash · ~17s/sample · retry · checkpoint Pipeline B: Qwen3-8B (3-Class)Gatekeeper→3-way Classify→MP→RewritevLLM · rules · dedup · 50K counting scans Pipeline C: Steerable (LLM-FREE)18-rule Filter→SAM3→700 Templates→Verify10K main + 2K reserve · verified Audit Pool → Balanced Subsets55,372 processed · 37,574 trainable · 3,815 early subset
Pipeline A

Gemini API: 4-Stage Agent — Real Sample Trace + Per-Stage Images

Sample: "Point to the sign." (source_rank=320928, gemini-3-flash-preview-thinking). Image+data merged per stage below.

Raw Recordquery, points_abs, label Gatekeeperkeep + core_target 5-Way ClassifyAff/Spatial/Reason/ObjRef/Count Multipointsingle vs multi Rewritecanonical format Processing workflow: cleaning→multipoint→classification_or_forced_counting→rewrite_same_session · status: completedFault tolerance: parse_retries=2, max_retries=6, sample_retries=4, ThreadPoolExecutor(32 workers)

Per-Stage Visualization — "Point to the DOG NOSE." (Object Reference → rewritten)

Stage 0

STAGE 0 — Raw Input

Query: "Point to the DOG NOSE."

Label: "DOG NOSE"

Point: (x, y) pixel coordinates

Source: pixmo_points dataset

Status: Unprocessed raw record — enters the 4-stage Agent pipeline

🔍 Orange circle with radar rings = search/scan context

Stage 1

STAGE 1 — Gatekeeper

Decision: KEEP (keep=true)

Core target: "the dog nose"

Rule 1: Must be a pointing task (GROUNDING) — ✓

Rule 2: Reject pure QA (text answers) — not applicable

Rule 3: Accept empty/free space pointing — not applicable

Rule 4: Accept logical pointing — not applicable

🟢 Green bounding box + "TARGET LOCKED" = validated and kept

Stage 2

STAGE 2 — 5-Way Classification

Category: Object Reference

Thought: "The task is a straightforward request to point to a specific named object without any spatial, functional, or quantitative constraints."

5 classes: Affordance | Spatial Relation | Reasoning | Object Reference | Counting

Method: Evaluates cognitive bottleneck, not surface keywords

Model: gemini-3-flash-preview-thinking · prompt_tokens=457

🔵 Blue point + 5 category chips (ObjRef highlighted) = classification result

Stage 3

STAGE 3 — Multipoint Detection

Result: single_point = TRUE

Reason: "The instruction refers to a singular specific object ('the dog nose') and does not request multiple points, spots, or all instances."

Multi triggers: all / several / each / every / multiple / free space

This sample: Unambiguous single target — no multi-instance language

🟣 Purple x1 marker + crossed-out xN = single-point confirmed

Stage 4

STAGE 4 — Rewrite

Output: "Point to the dog's nose — the prominent sensory organ on the animal's face."

Method: Category-specific REWRITE_PROMPTS["Object Reference"] applied

Effect: Preserves object identity while enriching with semantic description

Affordance: Would add interaction/usage cues ("where you would grip...")

Reasoning: Would add world-knowledge context ("the uppermost body part...")

🟢 Green text = rewritten output with category-specific enrichment applied

10 Balanced Samples — Skill-Based Rewrites (3 Aff + 3 Spatial + 1 Reason + 3 ObjRef)

s1
S1 [Affordance]: "Point to the place to hold coffee cup."
Point to the place for holding a coffee cup — the recessed area in the lid designed for st
s2
S2 [Affordance]: "Point to the area for cutting."
Point to the surface intended for cutting — the flat preparation area suitable for slicing
s3
S3 [Affordance]: "Point to the handle where you would hold the knife"
Point to the knife handle — the gripping part of the kitchen tool designed for safe holdin
s4
S4 [Spatial Rela]: "Point to the Bare area in front of rabbit left eye"
Point to the bare area located directly in front of the rabbit's left eye — the exposed sk
s5
S5 [Spatial Rela]: "Point to the yellow bird on top of head."
Point to the yellow bird positioned on top of the head — the small decorative figure situa
s6
S6 [Spatial Rela]: "Point to the rock."
Point to the rock situated on the ground surface — the geological object positioned below
s7
S7 [Reasoning]: "Point to the head."
Point to the head — the uppermost anatomical structure of the human body, identified as th
s8
S8 [Object Refer]: "Point to the window."
Point to the window — the specific text field on the form where a
s9
S9 [Object Refer]: "Point to the DOG NOSE."
Point to the dog's nose — the prominent sensory organ visible on the animal's face, identi
s10
S10 [Object Refer]: "Point to the sign."
Point to the sign — the specific information board shown in the image, recognizable by its

10-Sample Data Table — Rewrite Applied per Category Skill

#QueryLabelGatekeeperCategoryMultiRewrite (skill-based)
1Point to the place to hold coffee cup.place to hold coffeekeep=trueAffordancefalsePoint to the place for holding a coffee cup — the recessed area in the
2Point to the area for cutting.area for cuttingkeep=trueAffordancefalsePoint to the surface intended for cutting — the flat preparation area
3Point to the handle where you would hold the knife.handle where you woukeep=trueAffordancefalsePoint to the knife handle — the gripping part of the kitchen tool desi
4Point to the Bare area in front of rabbit left eye.Bare area in front okeep=trueSpatial RelationfalsePoint to the bare area located directly in front of the rabbit's left
5Point to the yellow bird on top of head.yellow bird on top okeep=trueSpatial RelationfalsePoint to the yellow bird positioned on top of the head — the small dec
6Point to the rock.rockkeep=trueSpatial RelationfalsePoint to the rock situated on the ground surface — the geological obje
7Point to the head.headkeep=trueReasoningfalsePoint to the head — the uppermost anatomical structure of the human bo
8Point to the window.windowkeep=trueObject ReferencefalsePoint to the window — the specific text field
9Point to the DOG NOSE.DOG NOSEkeep=trueObject ReferencefalsePoint to the dog's nose — the prominent sensory organ visible on the a
10Point to the sign.signkeep=trueObject ReferencefalsePoint to the sign — the specific information board shown in the image,

Rewrites follow category-specific skills from rewrite_skill.md: Affordance adds functional/interaction cues, Spatial preserves geometric relations, Reasoning adds world-knowledge context, Object Reference adds distinguishing features.

Pipeline B

Qwen3-8B Local: 3-Class + Rules + Dedup

Same 4-stage Agent via vLLM locally. 3 classes (Affordance/ObjRef/Reasoning). 18 keyword + 14 regex rules refine output. Dedicated Counting pipeline.

GatekeeperValidates + extracts target 3-Way ClassifyAffordance / ObjRef / Reasoning MultipointSingle vs multi-point RewriteCanonical format Rule Refinement (three_type_rules.py)18 affordance keywords + 14 reasoning regexrefine_three_type_category() Dedup (build_exclusion_keys.py)exclusion_image_keys.json (614K entries)nodup_reverse: filter B outputs already in Pipeline A Counting Pipeline (build_counting_direct.py)50K selected scans across four batchesSeparate track, Counting-specific rewrite Output: 3-class JSON + Counting JSONL + Dedup Map

10 Rule-Application Samples

#Input Query3-Class ResultRules TriggeredAction
1Point to the place to hold coffee cup.Affordance"place to hold" keywordKeep · Dedup check
2Point to where you would type in search bar.Affordance"where you would" + "type"Keep · Dedup check
3Point to the handle to open door.Affordance"handle to open" keywordKeep · Dedup check
4Point to all cups in the image.Counting → None"all" → multi-instanceRoute to Counting pipeline
5Point to the existing point + 5px right.Object Reference"existing point"/"cursor"Force ObjRef (special rule)
6Point to the direction the car is moving.Reasoning"direction" + "moving"Keep · reasoning_candidate_cue
7Point to the most likely exit.Reasoning"most likely"Keep · reasoning pattern
8Point to each person wearing a hat.Counting → None"each" → multi-instanceRoute to Counting pipeline
9Point to the wall.Object Reference— (no cue)Keep · cross-check vs Pipe A
10Point to the floor.Spatial → Nonespatial domainDrop (B keeps 3 classes)
Pipeline C

Steerable Pipeline: 5 Sub-Types (LLM-FREE)

5 task types with distinct pipeline logic. 700 templates x 7 groups, SAM3 masks, CC path verification. The main accepted set has 10K verified samples, with a 2K reserve batch.

Step 1: 18-Rule Filter + SAM3 Masks → Step 2: 700 Templates → Step 3: Path Verify Type 1 (85%)move_until_axis_fixedanchor→directiontarget on same axis8,500 samples Type 2 (5%)nearest_among_objectsanchor→directionnearest in path500 samples Type 3 (5%)nearest_axis_constraineddirection + axisconstraint500 samples Type 4 (2.5%)move_then_axis_limitedmove stepaxis-limited search250 samples Type 5 (2.5%)move_until_from_otheraxis from otherobject reference250 samples Step 4: select_one_candidate_per_image() + CC Path VerifyTiebreak: nearest > nearest_axis > move_then > move_until_other > move_until_axis_fixed OUTPUT: accepted_samples.jsonl — 10,000 verified samples across 5 task types{question, anchor, target, direction, task_type, image_id} · path verified · masks validated

SAM3 Mask Generation

Preview
SAM3 Preview: 8 objects from control points. Each color = one instance mask.
Overlay
Overlay on source image. CC path verification checks along directional arrow.

10 Steerable Samples — 2 per Sub-Type

Type 1: axis_fixed (85%)

T0
S1: "Use the blue point as your reference and point to the white stove to t"
anchor=[0.873,0.826] → target=[0.946,0.838]
T1
S2: "Take the blue point as the guide and find the quilt fabric to the righ"
anchor=[0.204,0.751] → target=[0.429,0.779]

Type 2: nearest_among (5%)

T2
S3: "Using the blue point only for distance, show the nearest buritto."
anchor=[0.128,0.443] → target=[0.184,0.311]
T3
S4: "Treat the blue point as your guide and locate the bud nearest to it."
anchor=[0.931,0.953] → target=[0.775,0.926]

Type 3: nearest_axis (5%)

T4
S5: "Treat the blue point as your guide and identify the nearest boulder to"
anchor=[0.484,0.510] → target=[0.342,0.485]
T5
S6: "From the blue point, identify the nearest gold square to the right of "
anchor=[0.803,0.453] → target=[0.936,0.454]

Type 4: then_axis (2.5%)

T6
S7: "Proceed up from the blue point, then turn left to select the dog."
anchor=[0.942,0.730] → target=[0.810,0.557]
T7
S8: "From the blue point, shift up, then shift right, to the hour hand."
anchor=[0.401,0.495] → target=[0.482,0.419]

Type 5: from_other (2.5%)

T8
S9: "Move diagonally up-left from the chestnut tree until you reach the gre"
anchor=[0.639,0.641] → target=[0.590,0.564]
T9
S10: "From the roof, move diagonally down-right until you reach the pool."
anchor=[0.722,0.068] → target=[0.883,0.381]

5 Sub-Type Pipeline Distribution

#Task TypeRatioSamplesPipeline Logic
1move_until_axis_fixed85%8,500anchor → straight direction → first target on same axis
2nearest_among_objects5%500anchor → direction scan → nearest object in path
3nearest_axis_constrained5%500anchor → direction + axis constraint → nearest in region
4move_then_axis_limited2.5%250anchor → move step → axis-limited target search
5move_until_from_other2.5%250anchor → axis from other object reference → target

10 Steerable Samples — Data Table

#Task TypeAnchorTargetQuestion
1axis_fixed[0.873,0.826][0.946,0.838]Use the blue point as your reference and point to the white stove
2axis_fixed[0.204,0.751][0.429,0.779]Take the blue point as the guide and find the quilt fabric to the
3nearest[0.128,0.443][0.184,0.311]Using the blue point only for distance, show the nearest buritto.
4nearest[0.931,0.953][0.775,0.926]Treat the blue point as your guide and locate the bud nearest to
5nearest_axis_constrain[0.484,0.510][0.342,0.485]Treat the blue point as your guide and identify the nearest bould
6nearest_axis_constrain[0.803,0.453][0.936,0.454]From the blue point, identify the nearest gold square to the righ
7then_axis[0.942,0.730][0.810,0.557]Proceed up from the blue point, then turn left to select the dog.
8then_axis[0.401,0.495][0.482,0.419]From the blue point, shift up, then shift right, to the hour hand
9from_other[0.639,0.641][0.590,0.564]Move diagonally up-left from the chestnut tree until you reach th
10from_other[0.722,0.068][0.883,0.381]From the roof, move diagonally down-right until you reach the poo
Innovation 2

AttnRes Architecture

Design

Gate-init=0 · Every 4 layers · LoRA r=64 + AttnRes (810 keys) · Dual-backbone (Steerability only)

Steerability Sweep

ConfigDefaultAttnResΔ
0-shot50.5%
n=259.5%63.5%+4.0
Standard Molmo2 vs AttnRes Architecture Standard Molmo2 Independent transformer blocks Block 1: Self-Attn + FFNBlock 2: Self-Attn + FFN Block 3: Self-Attn + FFNBlock 4: Self-Attn + FFN ... Block 11: Self-Attn + FFNBlock 12: Self-Attn + FFN No explicit cross-block history path AttnRes-Enhanced Molmo2 Every 4 layers: gated attention over historical block outputs Block 1Block 2 Block 3Block 4 ... Block 11Block 12 AttnRes Module 1 Attention over blocks 1-4 tanh(gate), gate_init=0 LoRA r=64 + saved AttnRes AttnRes Module 2: blocks 5-8 AttnRes Module 3 Attention over blocks 9-12 uses prior hidden states 810 trainable keys; dual-backbone routing for Steerability
Left: Standard Molmo2 keeps the normal top-to-bottom block flow. Right: AttnRes adds a tight side branch beside the enhanced blocks; the dashed history arrows stay next to the AttnRes modules and point right into the module edge.

Claim: AttnRes is a targeted fix for anchor-relative reasoning, not a universal backbone.

The clean architecture comparison is the n=2 setting: Default reaches 59.5% Steerability, while AttnRes reaches 63.5%. The same report also shows the boundary condition: too little Steerable signal underuses the module, while too much repetition overfits.

Ablation: architecture

Default n=259.5%
AttnRes n=263.5%

Same Steerable weight, different architecture. The +4.0pp gain supports the cross-block residual module.

Parameter sweep: data weight

n=1Default +2.5
n=2AttnRes +4.0
n=4Default +2.0

AttnRes needs enough anchor-relative data, but repeated Steerable samples hurt generalization.

Boundary: do not over-compose

Pure AttnRes63.5%
+Anchor enc.62.0%
+Iter corr.62.0%

Adding PointMLP+ViT on top of AttnRes does not improve peak Steerability, so the module should stay focused.

Attribution: Steerability instructions require propagating an anchor state through directional steps. Absolute point encoders help locate coordinates, but they do not replace cross-layer state flow. That is why ABC-C reaches 62.5% while pure AttnRes_n2 reaches 63.5%.
Detailed AttnRes Evidence Tables ablation + sweep

Steerability Sweep: When AttnRes Helps

Steerable WeightDefault BestDefault StepAttnRes BestAttnRes StepWinnerMeaning
n=1 (2,536 steer)62.5%2k60.0%18kDefault +2.5signal too small for extra module
n=2 (5,072 steer)59.5%10k63.5%14kAttnRes +4.0best cross-block regime
n=4 (10,000 steer)60.0%14k58.0%2kDefault +2.0repetition overfits

AttnRes + Anchor Encoding Ablation

ExperimentArchitectureExtra Point EncodingCorrection TrainingBest SteerBest Ckpt/RoundTakeaway
AttnRes_n2 sweepAttnResNoneNo63.5%ckpt-14000pure AttnRes LoRA is strongest
steer_baseAttnResPointMLP + ViTNo62.0%ckpt-16000 R0anchor encoding does not add gain
steer_iterAttnResPointMLP + ViTYes62.0%ckpt-2000 R2iter helps R2 stability, not peak
ABC-C referenceDefaultPointMLP + ViT + CoordMapYes62.5%ckpt-2000 R2encoding helps, but below AttnRes_n2

Why Dual Backbone Is Necessary

ModelAfford.CountingReason.SpatialSteer.Interpretation
steer_base ckpt-16000 R091.9%16.3%76.7%77.9%62.0%Steerable-only recipe destroys Counting
steer_iter ckpt-2000 R287.4%23.5%76.7%77.9%62.0%correction recovers rounds, not category balance
Innovation 3

ABC Encoding

ABC
A: text-only correction. B: PointMLP + ViT visual features. C: B plus CoordMap CNN.

Claim: the main gain comes from grounding perturbed coordinates in visual features.

A text-only correction learns the wording but cannot reliably bind the perturbed point to image evidence. B injects a learned point embedding plus ViT features and gives the core gain. C adds a coordinate map, which mostly helps Spatial but has lower ROI.

Ablation: text-only is insufficient

A overall72.3%
B overall75.5%
Gain+3.2pp

Same correction task, but B injects point + visual features instead of relying on text coordinates.

Where B helps most

Counting+7.2pp
Reasoning+7.3pp
Afford.+3.0pp

Point-conditioned visual features help multi-point perception and spatial-semantic reasoning.

C is useful, but targeted

C overall77.0%
C vs B+1.5pp
Spatial+3.6pp

CoordMap mainly contributes dense spatial context; Reasoning is slightly better with B.

Core ABC Ablation

ModulePoint RepresentationOverallCountingReasoningSpatialInterpretation
Atext-only coordinates72.3%61.7%72.5%80.0%weak binding between text coords and image features
BPointMLP + ViT features75.5%68.9%79.8%79.0%best cost/performance point encoding
CB + CoordMap CNN77.0%69.9%79.3%82.6%best final score; Spatial-specific gain
Attribution: the loss report shows fixed protocol tokens and label text are already learned, while coordinate digits dominate 91-95% of the supervised CE. The model needs a representation that aligns coordinates with visual evidence, not more text-only coordinate formatting.
Detailed ABC Evidence and Failure Analysis ablation + attribution

Correction Rounds: Iteration Is Not the Main Source of Gain

ArchitectureR0 BestR1 BestR2 BestBest RoundObserved Effect
A Text-only72.3%72.2%72.0%R0late training improves text-only but remains below B/C
B PointMLP+ViT75.5%75.5%75.5%R0/R1/R2 tiepoint encoding already resolves most correction signal
C +CoordMap77.0%75.5%76.7%R0CoordMap is strongest early; extra rounds slightly decay

B vs C ROI

MetricB BestC BestC Gain
Overall75.5%77.0%+1.5
Counting68.9%69.9%+1.0
Reasoning79.8%79.3%-0.5
Spatial79.0%82.6%+3.6
Steerability61.5%62.5%+1.0

Training Objective Diagnosis

SignalStep 2000Step 20000What It Shows
Total CE0.62350.5474loss keeps falling
Overall eval73.42%73.01%metric does not rise
Coord loss share91.19%95.38%coordinates dominate errors
Label loss0.54250.1757text label already learned
Coord ones loss2.13592.0295fine digits remain hard

SinglePOS Negative Control

EncodingFormatBest OverallCountingStatusLesson
Baseline DMolmo2 HTML continuous coords73.42%67.3%stablekeep native coordinate protocol
singlepos originalN POS_x POS_y...64.56%40.8%learns slowlydiscrete tokens need strong structure
singlepos_longPOS_x POS_y...13.95%1.5%failedno count prefix breaks structure
legacy_long1 x y 2 x y...56.11% -> 13.24%32.1% -> 0.5%epoch-2 collapsesequence can loop without total count
count_longN 1 x y 2 x y...14.66%2.6%failedextra structure dilutes POS learning
Final Results

Multi-Expert Routing

Claim: the final score is the end of a proof chain, not a single checkpoint.

The reports break into three evidence types. Ablations show which module removes which failure mode. Parameter sweeps only tune how much signal each module gets. Negative controls reject simpler substitutes. The final ensemble wins because it routes each category to the expert that matches its error pattern.

Evidence map: `ABC_analysis.md` and `steerability_sweep_report.md` are module evidence; `runs_overview_report.md` is hyperparameter tuning; `metadata_prompt_variants_A_F_report.md` and `singlepos_experiments_report.md` are negative controls; `output_D_lr2e4_ga16_training_eval_loss_report.md` and `output_D_lr2e4_ga16_counting_error_analysis.md` explain why the gains do or do not transfer.
Step 1 · data
55,372 → 37,574

Audited pools feed balanced runs

The cloud inventory has 53,772 unique sample IDs and 37,574 trainable rows. The 3,815-row set is only the early semantic baseline subset.

Step 2 · ablation
72.3 → 75.5 → 77.0

ABC fixes coordinate-image binding

Text-only A underperforms. B gets the core gain from PointMLP + ViT. C mostly buys Spatial, so B is the better ROI while C is the higher ceiling.

Step 3 · tuning
59.5 → 63.5

AttnRes only helps at the right weight

n=2 is the sweet spot. n=1 underuses the module and n=4 overfits. That makes AttnRes a targeted fix for anchor-relative Steerability, not a universal backbone.

Step 4 · routing
77.2%

Category experts win

Affordance C 93.94, Counting C 70.41, Reasoning B 78.24, Spatial C 82.56, Steerable AttnRes 63.00. These are local expert-selection scores; the final overall is the submitted benchmark score.

Proof Table 1: Data Base Before Model Changes

Funnel view: cloud inventory first, experiment subsets second

Processed outputs
55,372
Unique IDs
53,772
Trainable rows
37,574
Mixed D subset
12,680
Early subset
3,815

The 3,815-row bar is a deliberate early semantic subset, not the total data volume. Later D-format and steerability runs resample larger balanced subsets from the audited cloud pools.

EvidenceCountOperationMarkAnalysis
Processed cloud outputs55,372merge Gemini + Qwen + SAM-clean batchescloud poolshows the full multi-batch inventory, not only the early filtered-data run
Unique sample IDs53,772deduplicate repeated workflowsclean baseremoves repeated attempts before counting usable samples
Trainable completed rows37,574latest judgement completed/acceptedmodule outputusable rows available before experiment-specific balancing
Early Gemini trainable rows24,415semantic cache onlyearly poolvalid for the original Pipeline A baseline, not the full server count
Early FT subset3,815763 per category x 5early balancesmall controlled subset used for the early semantic comparison

Attribution: the data module is not just more rows. The full cloud pool provides 37,574 trainable rows; the 3,815-row set is a deliberately balanced early subset where Spatial's 763 eligible examples define the quota.

Proof Table 2: Module Ablation, Not Hyperparameter Tuning

Overall ablation: A to B is the core jump

A text-only
72.3%
B PointMLP+ViT
75.5%
C + CoordMap
77.0%

B vs A is the module proof (+3.2pp). C vs B is a smaller, targeted spatial refinement (+1.5pp).

Where the gain comes from

Counting B-A
+7.2
Reasoning B-A
+7.3
Spatial C-B
+3.6
Overall C-B
+1.5

The biggest jumps are Counting and Reasoning from visual point grounding; CoordMap mainly helps Spatial.

ModuleRepresentationOverallAfford.CountingReason.SpatialSteer.MarkAnalysis
Atext-only coordinates72.3%90.9%61.7%72.5%80.0%61.5%baselinewording can be learned, but coordinate-image binding is weak
BPointMLP + ViT75.5%93.9%68.9%79.8%79.0%61.5%core ablationmain gain comes from point-conditioned visual grounding
CB + CoordMap CNN77.0%94.4%69.9%79.3%82.6%62.5%targeted boostbest final score, but its distinctive value is Spatial

Attribution: B over A is the strongest module proof: Counting +7.2pp and Reasoning +7.3pp indicate that visual point features solve a grounding problem. C is still useful, but its extra value is concentrated in Spatial (+3.6pp), so the page should describe C as targeted rather than universally better.

Proof Table 3: AttnRes Architecture and Steerable Weight Sweep

Default vs AttnRes by steerable weight

n=1 Default
62.5
n=1 AttnRes
60.0
n=2 Default
59.5
n=2 AttnRes
63.5
n=4 Default
60.0
n=4 AttnRes
58.0

Only n=2 is the clean architecture proof: same data weight, AttnRes +4.0pp over default.

Interpretation of the sweep

n=1 signal
too low
n=2 signal
sweet spot
n=4 repeat
overfit
Peak step
14k

The sweep is not just ranking models; it explains the usage rule: enough steerable data for AttnRes, but not repeated to saturation.

Steerable WeightDefault BestDefault StepAttnRes BestAttnRes StepWinnerEvidence TypeInterpretation
n=1 / 2,536 steer62.5%2k60.0%18kDefault +2.5boundarytoo little steerable signal underuses AttnRes
n=2 / 5,072 steer59.5%10k63.5%14kAttnRes +4.0architecture proofsame data weight, cross-block residual attention wins
n=4 / 10,000 steer60.0%14k58.0%2kDefault +2.0overfit boundaryrepetition hurts generalization and narrows the useful regime

Attribution: AttnRes should be claimed narrowly. The n=2 row proves the architecture helps anchor-relative Steerability; n=1 and n=4 explain why it must be paired with the right data weight and early stopping.

Proof Table 4: Negative Controls and Failure Attribution

Rejected shortcuts

No meta prompt
71.38
Best prompt D
71.69
SinglePOS best
64.56
SinglePOS fail
14.66

Prompt-only and token-format shortcuts do not explain the final gain; most remain below the trained modules.

Counting error anatomy @ ckpt-2000

Success
132/196
Wrong count
50/196
Out of mask
44/196
Invalid parse
5/196

Counting fails structurally: wrong number of points and mask misses dominate; parser failures are small.

TestBest / Key ResultCountingSteer.MarkWhat It Rules OutWhat It Proves
Prompt metadata A-FD = 71.69%, C = 71.38%66.84%50.5%small deltaprompt-only coordinate injection is not enoughtraining modules are needed, not just coordinate wording
SinglePOS original64.56%40.8%53.0%negative controldiscrete POS tokens do not solve coordinates by themselvesoutput structure and coordinate grounding remain bottlenecks
SinglePOS failed variants13.95% / 14.66%1.5% / 2.6%12.5% / 13.0%collapseformat changes can collapse parsing and point countnative Molmo2 coordinate protocol is safer
Loss diagnosisCE 0.6235 -> 0.5474, eval 73.42 -> 73.01loss highest 0.6200label loss near 0metric gaplower token CE does not imply better mask accuracycoordinate digits dominate 91-95% of supervised CE
Counting error anatomy132/196 success50 wrong count, 44 out maskN/Arouting causeerrors are not mainly parser failuresCounting needs a category expert for multi-point structure

Attribution: the negative controls prevent overclaiming. The method advantage is not from a lucky prompt or a coordinate-token trick; it comes from modules that match the real failure modes.

Data module

0-shot72.7%
Pipeline A75.3%
Trainable37,574

LLM cleaning, local filtering, steerable synthesis, and balanced resampling lift semantic categories before model specialization.

ABC module

A overall72.3%
B overall75.5%
C overall77.0%

Point-conditioned visual features recover coordinate grounding. B is the best ROI point; C is the best Spatial point.

Negative controls

Prompt D71.69%
SinglePOS best64.56%
Count valid fmt82.65%

Prompt-only changes buy little. Discrete POS variants mostly fail. The 82.65% Counting number is valid-output rate, not mask accuracy.

Final Routing Chart: Category Experts

Per-category expert accuracy

Affordance · C
93.94%
Counting · C
70.41%
Reasoning · B
78.24%
Spatial · C
82.56%
Steerable · AttnRes
63.00%

These category bars are local validation evidence for choosing experts. The official submitted overall score is reported separately as 77.2%.

#CategoryExpertAccuracyWhy This ExpertMark
1AffordanceC93.94%C keeps the highest semantic grounding scoreroute
2CountingC70.41%multi-point structure benefits from visual + CoordMap groundingroute
3ReasoningB78.24%B has better ROI and avoids C's slight Reasoning dropB wins
4SpatialC82.56%CoordMap's clearest gain is Spatial +3.6pp over BC wins
5SteerableAttnRes63.00%anchor-relative reasoning needs cross-block state flowAttnRes

Official submitted overall: 77.2%

Category values above are local validation scores used for routing, not a re-summed 982-sample ensemble.

Complete Evolution: What Each Stage Adds

Overall progression

Molmo2 0-shot
72.7%
+ Pipeline A
75.3%
+ IterCorr B
75.5%
Final Ensemble
77.2%

The final gain is not monotonic in every category. Pipeline A improves semantic categories but hurts Counting/Steer; later modules recover hard categories, then routing selects per-category experts.

ExperimentAfford.SpatialReason.Count.Steer.OverallAnalysis
Molmo2 (0-shot)85.9%76.9%77.2%73.0%50.5%72.7%strong base, but Steerability is the first clear bottleneck
+ Pipeline A93.9%83.1%82.9%67.3%49.5%75.3%data synthesis improves semantic pointing but does not solve hard structured categories
+ IterCorr B93.9%79.0%79.8%68.9%61.5%75.5%point-conditioned visual grounding recovers Steerability and improves coordinate binding
Final Submission93.9%82.6%78.2%70.4%63.0%77.2%official overall score; category cells show local expert-selection evidence
Why routing wins: Counting errors are mostly wrong counts and out-of-mask points; Steerability needs cross-block anchor state; Spatial benefits most from CoordMap. The expert map is therefore a consequence of the failure analysis, not a cosmetic ensemble trick.
Detailed Diagnostics by Evidence Type supporting evidence
Read these tables as support for the proof chain: cloud-pool auditing and balancing explain the training base; LR/GA and checkpoint timing are parameter sweeps; prompt-only and SinglePOS are negative controls; Counting error tables are failure attribution for category experts.

Cloud Data Inventory Behind the Training Sets

StageCountFilter / ModuleRole
Processed outputs55,372Gemini + Qwen + SAM-clean cachefull audited cloud inventory
Unique sample IDs53,772latest-workflow selectiondeduplicate repeated processing roots
Trainable completed rows37,574completed / accepted latest judgementusable pool before experiment balancing
Early Gemini cache24,415semantic LLM pass onlysource for the original filtered-data baseline
Early balanced subset3,815763 per category x 5controlled semantic LoRA comparison

Category Balancing: Why the Selector Matters

CategoryEligibleSelectedDroppedWhy It Matters
Affordance1,7787631,015avoid overfitting easy category
Counting12,89576312,132massive raw skew needs quota control
Object Reference4,2247633,461keeps semantic pointing balanced
Reasoning2,4787631,715preserve harder cognitive tasks
Spatial Relation7637630scarce class defines the quota

Final source mix: 3,715 pixmo_points + 100 where2place. Robopoint was abundant but not selected after balancing.

Filtered-Data FT Grid: Best Runs

RunLR / GABest StepOverallAfford.SpatialReason.Steer.CountingInvalid
lr1e4_ga161e-4 / 16200075.25%93.94%83.08%82.90%49.5%67.35%9
lr2e4_ga162e-4 / 16200074.85%90.91%81.54%78.24%54.0%69.90%4
lr1e4_ga81e-4 / 8200073.83%93.43%83.08%76.17%52.0%64.80%7
lr2e4_ga82e-4 / 8200014.87%31.82%16.92%10.88%12.5%2.04%103

Higher LR with GA=8 collapses output format; stable LoRA needs the larger effective batch.

Prompt-Only Metadata Injection: Limited Gain

VariantCoordinate PromptOverallSteer.InvalidConclusion
Anatural-language pixel coords69.55%40.5%1extra wording hurts Steerability
Bnatural language + Molmo2 HTML69.04%38.0%1format match alone is insufficient
Cnone71.38%49.5%1baseline prompt
Ddirect legacy coordinate list71.69%50.5%3only prompt variant above no-injection
Ebare HTML points tag68.33%34.5%2empty tag disrupts generation
Fnormalized bracket list70.57%45.5%1still below D/no-injection

Prompt engineering cannot replace the training modules; direct D format is the only useful metadata prompt.

Counting Error Anatomy

Metric @ ckpt-2000Value
Counting samples196
Success132 / 196 = 67.35%
Wrong number of points50 / 196 = 25.51%
At least one point out of mask44 / 196 = 22.45%
Correct count but out of mask14 / 146 = 9.59%
Wrong count and out of mask30 / 50 = 60.00%
Invalid parse5 / 196 = 2.55%

Counting failures are mostly structured multi-point mistakes, not only parsing errors.

Counting-Specialized Model

SettingValue
ArchitectureAttnRes + C wrapper
BalanceCounting x2, others x1
Train data18,516 samples
Best checkpointckpt-1000
Counting valid output162 / 196 = 82.65%
Wrapper paramsPointMLP + ViT + CoordMap CNN
Use in storyformat diagnostic, not accuracy

Source: model_handover_AttnRes_n2.md and SUBMISSION_REPORT.md. The 82.65% value is valid-output rate; its Counting accuracy was lower than the C route's 70.41%.

Checkpoint Timing: Early Stop Beats Longer Training

Run / ModelEarly BestLate ScoreLoss TrendLesson
lr1e4_ga1675.25%@2k69.65%@16ktrain loss 0.553 -> 0.311eval overfits while CE falls
output_D_lr2e4_ga1673.42%@2k73.01%@20kCE 0.6235 -> 0.5474digit loss drop does not equal mask accuracy
ABC-C77.0%@2k R073.8%@20k R0extra CoordMap peaks earlyselect early checkpoint for final expert
AttnRes_n263.5%@14k60.0%@20kpost-peak declineSteerability needs category-specific early stop