RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026Start here · 100 retrospective records ↗
AI Use Field Guide

Start here / Prompting & checks

Prompting & checks / Start-here guide · Start-here guide · prepared 16 September 2026

Telling a model to act as an expert rarely lifts accuracy

A 2023 study of 162 personas found no general accuracy gain, while a vendor guide frames roles as a tone setting.

platform.claude.comprimary record

Prompting best practices

Document
undated document
Event
no single event
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

Start here

Opening a prompt with 'you are a senior tax attorney' feels like it should sharpen a model's answer, and it is common enough advice to seem settled. A specific study of that claim, and a current vendor guide's own description of what a role does, both point to a narrower effect than the advice usually implies.

What the documents say

Anthropic's current prompting reference describes giving a role in the system prompt as something that focuses Claude's behaviour and tone, adding that even a single sentence makes a difference, a claim about tone and focus, not about factual accuracy. The direct test of the accuracy claim is When 'A Helpful Assistant' Is Not Really Helpful, submitted 16 November 2023 by Mingqian Zheng and four co-authors. It is a systematic evaluation: 162 personas covering six types of interpersonal relationship and eight domains of expertise, tested across four model families on 2,410 factual questions. Its finding is direct: adding personas in system prompts does not improve model performance across a range of questions compared to a no-persona control. A second layer is worth keeping separate from the headline: a persona chosen specifically for one question can raise that question's accuracy, and gender, relationship type and domain all show some influence, but automatically picking the best persona in advance performed no better than random selection. This is a study of accuracy on factual questions, not of tone or reader preference, which the paper does not measure.

Check this

The check that separates the two effects: give a model the same factual question with and without an expert persona, several times each, and compare correctness, not confidence or fluency. Separately, compare the same pair for tone or format only, on a task with no single correct answer, such as a draft email. The vendor's own framing predicts a difference in the second test and the study predicts little difference, on average, in the first.

What holds and what fails

A role instruction holds up, per the vendor's own description, as a lever on tone and behavioural focus. It fails as a general accuracy booster for factual questions, per the study's controlled comparison against a no-persona baseline, even though a well-matched persona can help a specific question and a poorly matched one can hurt. Generalising either result past its own measure, treating a tone effect as an accuracy effect, or one favourable persona as proof the technique reliably improves correctness, is the overreach both documents together rule out.

  • Use a role instruction to set tone and audience, not as a substitute for checking a factual answer.
  • If accuracy on a factual task matters, test the same question with no persona at all as a baseline.
  • Do not assume a persona that helped once will help on a differently worded version of the same question.

The two documents are not in conflict: one describes what a role is built to do, and the other measured what it does not reliably do, and keeping that boundary in view is more useful than treating role prompts as either useless or essential.

Sources & reading trail

Prompting best practices ↗

States that a system-prompt role focuses behaviour and tone, and that even one sentence makes a difference.

Source published: Not established · Retrieved: 16 September 2026

When ‘A Helpful Assistant’ Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models ↗

Reports the 162-persona, 4-model-family, 2,410-question evaluation finding no general accuracy gain from personas.

Source published: 16 November 2023 · Retrieved: 16 September 2026

Documentation, regulator guidance and studies establish the record; the checks and the boundary are AI Use Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.