Only the disagreements count: McNemar’s test for robot policies
You have two policies for a pick-and-place task. Maybe one is a fine-tuned checkpoint and the other is the baseline, or maybe one has an extra camera. You run each of them 50 times. Policy A succeeds 42 times (84%) and policy B succeeds 34 times (68%). Sixteen percentage points...

