Figure 1: Results on RULER and OOLONG-synth. Our finetune runs in our harness with an 8K token
per-agent context limit. GPT-5.4 (medium reasoning) and base Qwen3.6-35B-A3B (thinking disabled) receive the full document with the prompt (see Section 4 and Appendix A for more details). Base Qwen and GPT-5.4 are served with 64K and 1M token context limits, respectively. The OOLONG random
baseline follows Bertsch et al. (2025, §2.3) .