Structure is Cheap, Grounding is Not
While fine-tuning enables 103M–135M parameter language models to effortlessly achieve near-perfect JSON structure when calling tools, they consistently fail to accurately bind arguments to contextual facts. This negative result demonstrates that while structural fluency is easy for small models to acquire, true state grounding remains a fundamental bottleneck at this scale.