Harness-1 edges GPT-5.4 on search recall by externalizing state, not scale
Harness-1, a 20-billion-parameter open-source search agent, scored 73% on average recall across eight search benchmarks, edging GPT-5.4's 70.9%. The 2.1-point margin is narrower than the source's "massive leap" framing suggests, and the comparison carries a