Giving an AI agent a detailed professional identity may sound like an inexpensive way to improve its expertise, but a new preprint finds little evidence that broad, profession-specific instructions improve answer accuracy. The paper, written by Timothy Kassis of K-Dense, evaluates a public collection of 503 scientific and engineering profiles designed for use as persistent system prompts.

The study tested profiles that describe a profession's workflows, analytical tools, common errors and reporting standards. Kassis compared a matched full profile with four alternatives: a minimal helpful-assistant prompt, the profile's one-sentence role description, a generic scientific-rigor guide and a similarly sized profile from an unrelated field. The user question stayed the same while only the system prompt changed.

Across nine text-based science tasks, the experiment sampled 4,531 questions and retained 4,488 items that completed all five prompt conditions after technical retries. The paper reports an average profile-versus-baseline accuracy difference of negative 0.6 percentage points, with a 95 percent bootstrap interval from negative 1.5 to positive 0.2 points. None of the nine benchmarks showed a statistically clear improvement from the matched professional profile.

AI-generated editorial illustration of five system-prompt conditions in a study
AI-generated illustration: Five system-prompt conditions in a study.

The full prompts were substantially more expensive to run. According to the paper, matched profiles generated 1.5 to 2.3 times as many output tokens and raised estimated list-price cost per successful call by 2.2 to 4.5 times compared with the minimal baseline. The long generic and mismatched prompts produced similar cost increases, suggesting that prompt length, not specialized scientific content, drove much of the added usage.

Performance was worse in a separate tool-using bioinformatics evaluation under fixed resource limits. On 60 BioMysteryBench problems, each run three times for the baseline and matched profile, the profile condition solved an average of 46.7 percent while the baseline solved 56.7 percent. The paper attributes much of that 10-point gap to more frequent token- and time-limit stops: the profile was resent on every turn, consuming more of the available budget.

One result cut in the opposite direction. On the first, unretried SuperGPQA pass, the short baseline suffered many more provider API failures. It returned a correct first-pass answer on 54.0 percent of scheduled items, compared with 71.6 percent for the matched profile. But generic and unrelated long prompts performed about as well as the matched profile, so the author argues that domain expertise was not responsible; the mechanism may instead involve prompt length or formatting, and the paper says it remains unclear.

AI-generated editorial illustration of higher token use without a clear accuracy gain
AI-generated illustration: Higher token use without a clear accuracy gain.

The study's scope limits how broadly the findings can be applied. Every main experiment used Gemini 3.8 Flash through OpenRouter in the Pi agent harness. Text answers were graded automatically with rule-based checks, which can score objective correctness but not qualities such as experimental design, literature synthesis or explanatory clarity. The results therefore do not establish that professional profiles are useless for open-ended scientific work.

The tool-using comparison also placed longer prompts at a structural disadvantage because each execution had a token cap and a 35-minute time limit. The author notes that incomplete logs prevent ruling out effects from reused cloud containers. The paper also says the scientific accuracy of all 503 profiles was not audited and benchmark overlap with model training data was not checked.

A further caveat is institutional: Kassis is affiliated with K-Dense, the organization that created the Scientific Agents collection being evaluated. The profile corpus is public, but the paper says its evaluation code and item-level records have not been released, limiting independent reproduction. The manuscript is an arXiv preprint rather than a reported peer-reviewed result.

Within those boundaries, the finding is a practical warning against loading large role manuals into every request by default. The study suggests teams should measure completed-answer accuracy, first-pass reliability and total resource use separately. It identifies selective retrieval—supplying only the relevant profile sections for a given task—as an unanswered next step that could preserve useful guidance without paying the full cost of persistent, profession-wide context.