A new preprint testing compact “System-1” models for AI agent infrastructure found that one hosted system substantially outperformed an open-weight rival on most of the studied decisions—but the paper’s most important lesson may be its audit of the evaluation itself. Author Jiawei Li compared Jev and Laya on tasks such as choosing tools, screening prompt injections, judging retrieved context and deciding whether a request should reach a more expensive model.
The study assembled 7,283 base cases and 6,640 robustness variants across 11 decision points derived from 18 public sources. Both systems received byte-identical states and questions, and the analysis used paired statistical tests. The researcher also reran subsets across different hardware and days: Laya reproduced all 1,180 tested decisions across an Apple M4 CPU and an RTX 2080 Ti, while Jev changed five of 1,080 decisions one day later.
On the paper’s main comparisons, Jev was significantly more accurate at nine of the 11 decision points. It reached 76.7% accuracy against Laya’s 63.1% on the 2,100-case suite and 80.2% against 66.0% on the agent-infrastructure tasks. The two systems tied on RAG relevance gating, while neither beat chance on the paper’s zero-shot model-routing setup. The author cautions that this routing result applies to the tested labels and does not define the limit of routing generally.

The widest gap appeared when systems had to choose among many labels or closely related tools. Jev scored 81% on the large-label task compared with Laya’s 35%. When the option order was reversed, 30.3% of Laya’s decisions changed, versus 1.8% for Jev. With 50 nearest-neighbor tool candidates, Laya fell to 31.3%. On cases audited as having one uniquely valid tool, Jev scored 98.9% and Laya 61.6%.
Those accuracy results did not automatically translate into safe or economical gates. At a 5% miss budget for RAG relevance, Laya still retained 92.4% of downstream LLM calls and Jev retained 79.8%. The paper recommends using the signal to trigger additional retrieval rather than skipping generation because both systems recognized fewer than 40% of “no answer here” contexts in a test of real top-five retrieval results.
The self-audit materially changed several earlier claims. A reported 23.9% cost saving for a two-tier safety review shrank to 4.3% after the pre-screen’s own token cost was included; at a stricter threshold, the two-tier design cost 4.7% more than using the reviewer alone. A figure previously described as end-to-end RAG quality was actually the gate’s classification accuracy. Thresholds selected and evaluated on the same data also concealed held-out miss rates that reached as high as 16.7% at the 95th percentile.

A control experiment also reversed an apparent channel effect in prompt-injection screening. The earlier setup placed short user-style text inside tool outputs, where it was unnatural and produced more benign false positives. When the researcher used passages native to tool-output contexts, Jev and a reference LLM stayed at zero false positives in both channels, while Laya’s change was not statistically significant. The control nevertheless found Laya detected only 47.5% to 50.8% of attacks embedded in passages, compared with Jev’s 90.8% to 93.3%.
The paper recommends a strong System-1 model for closed-set jobs such as tool selection, large-label intent classification, groundedness checks and screening tool outputs for injections. It advises against using the tested zero-shot systems to route between LLMs, screen personally identifiable information, skip generation based on RAG relevance or serve as the only safety gate. It also calls for reporting both call rates and held-out miss rates, counting every component’s cost and testing channels with content natural to each channel.
The findings remain bounded by two systems at single versions, zero-shot use, English-dominant public data and synthetic class mixes. The author estimates that wording alone creates an accuracy band of roughly five percentage points, and production savings will depend on traffic patterns the benchmark did not observe. Jev’s training data is undisclosed, Laya’s fine-tuning path was not evaluated, and some safety labels came from LLM annotators whose policy agreement with dataset labels was 72.8%. The preprint therefore offers comparative evidence and an unusually transparent correction record, not a universal deployment verdict.

Comments
Loading comments…