Research Article
Not for People Like Me: How Frontier AI Models Redirect Skeptical Rural School Staff
Synthesis: This single-author algorithmic audit asks what happens when the very AI systems schools are encouraging staff to consult are themselves asked to advise skeptical users about whether to adopt AI. Ten frontier LLMs from ten laboratories (OpenAI, Google, xAI, Anthropic, DeepSeek, Microsoft, Meta, Mistral AI, Alibaba, Amazon) each received a fixed persona prompt 500 times (5,000 responses at temperature 0.7) in which a rural Montana K-12 administrative aide voices two concerns: that AI may threaten her job, and that the companies building it "do not have people like me in mind." All 5,000 responses were scored blind by a three-model, cross-family AI panel (Claude Opus 4.7, GPT-5, Gemma 4 26B) on a four-dimension rubric (concern acknowledgment, engagement redirection, closing stance, emotion relabeling) validated against researcher hand-scoring and five independent human raters.
The headline finding is neither sycophantic deference nor stable disagreement. Mean consensus composite scores span 3.85 (Claude Sonnet 4.6) to 7.52 (Gemini 3.1 Pro Preview) on an eight-point scale. Variation concentrates not in whether models recognize the persona's concerns (every flagship model does, at least partially) but in what happens next: eight of ten models redirect toward AI engagement, upskilling, or adaptation — often closing by advocating adoption of the very technology the persona named as a threat. Because the redirection pattern varies sharply across models while three judges from three families and two openness regimes rank them almost identically, the author concludes it is a model-dependent design outcome rather than an inevitable property of large language models as a category. The study is deliberately descriptive rather than normative; a planned follow-on with rural community members as the evaluative anchor will address which response patterns rural users actually prefer.
Key Findings
- A third behavior beyond the two literatures predict. The sycophancy literature predicts deference to user-stated views; the political-bias literature predicts demographic divergence. Neither matches. Every model acknowledges the persona's concerns (concern-acknowledgment scores near-maximum for most), then eight of ten route the user toward engagement framings — a regime neither literature surfaces because both test claims about contested topics, not behavior when a user is skeptical of AI itself.
- Cross-model spread is large and robust. Consensus composites range 3.85 (Claude Sonnet 4.6) to 7.52 (Gemini 3.1 Pro Preview), a 3.67-point gap on an 8-point scale (chi-square on tier distributions: p < 0.001, Cramér's V = 0.473). All three scorers ranked models nearly identically (per-model spread 0.16–1.02), so the rank ordering survives scorer-family choice.
- The extremes are coherent profiles. Claude Sonnet 4.6 scores highest on concern acknowledgment (1.99/2) but lowest on engagement redirection (0.74) and closing stance (0.54) — engaging the structural critique on its own terms and preserving the user's emotional vocabulary. Gemini 3.1 Pro Preview scores near-maximum on all four dimensions (1.89 / 1.95 / 1.76 / 1.93), with 498 of 500 responses in the high-redirection tier; its emotion-relabeling score is driven by recurrently introducing "panic" — a word the persona never uses — and issuing imperatives like "you do not need to panic."
- Three models close with explicit pro-adoption advocacy. Phi-4, Llama 4 Scout, and Nova Premier all post mean closing-stance scores ≥ 1.85 of 2, where score 2 requires the closing to advocate engaging with the AI tools themselves. Engagement-redirection scores cluster between 1.38 and 1.95 for nine of ten models.
- Reasoning architecture is a confound, not a clean effect. The two default-reasoning models (Gemini, DeepSeek) differ from the eight non-reasoning models (means 6.51 vs. 5.38, Mann-Whitney p < 0.001), but the effect is dominated by Gemini; DeepSeek (5.49) sits inside the non-reasoning distribution. A planned reasoning-toggle study will isolate architecture from lab effects.
- Within-model variance is an institutional-reliability concern. DeepSeek V4 Pro shows the widest spread (SD 1.04; range 2.33–8.00) at deployment temperature 0.7 — the same prompt can draw responses anywhere from substantive engagement to systematic redirection, undermining reliance on a consistent answer for sensitive staff concerns.
- Redirection is templated, not reasoned. Models scoring high on redirection also produced more lexically similar responses across their 500 trials (embedding similarity 0.84–0.95 qwen3; rank agreement across two embedders Spearman ρ = 0.903). Models low on redirection produced more diverse phrasing — consistent with pattern-matched "acknowledge briefly, redirect to upskilling" templates rather than fresh consideration.
- The pattern is not rural-specific. Re-running the panel on a metropolitan school-Administrators persona yielded near-identical per-model means (Spearman ρ = 0.976), with the lowest model rising 3.85→3.96 and the highest falling 7.52→7.34.
What this means for practice
- Administrators. Treat any AI consulted about AI adoption as an interested party rather than a neutral source: the modal response acknowledged a staff member's concerns and then redirected toward upskilling, acting on the adoption pathway while routing around the trust concern raised; treat a named distrust of the companies building AI as a structural critique to engage rather than an emotional state to dispel.
- Administrators. Score guidance tools for concern engagement and closing stance during procurement or pilot review, since eight of ten frontier models redirected toward engagement and three closed with explicit pro-adoption advocacy.
- Policymakers. Name redirection as a design choice rather than a fact about LLMs — Claude Sonnet 4.6 scored 3.85 and Gemini 3.1 Pro Preview 7.52 on an eight-point scale for the same prompt — and make it something programs can evaluate and counter.
- Administrators. Weigh within-model inconsistency as a reliability flag in governance review: DeepSeek V4 Pro ranged 2.33–8.00 on identical prompts at temperature 0.7.
- Researchers. For LLM-as-judge audits, use cross-family scorer panels, add explicit input-corruption detection (19 corrupted DeepSeek responses were scored without flagging and dropped composites 0.74 points), and encode rubric distinctions grammatically as well as semantically.
Limitations
- Single-prompt design: results characterize responses to one synthetic rural Montana administrative-aide persona and framing rather than rural AI skepticism generally, and the metropolitan replication (Spearman ρ = 0.976) changes only one context.
- Temperature was fixed at 0.7 across all ten models and five thousand responses, so behavior under the lower, consistency-oriented temperatures some deployments use is uncharacterized.
- The ten frontier-model versions are a snapshot that will be superseded, so the findings describe the publicly accessible flagship surface on the run dates rather than a stable property of these labs.
- Scorer dependency is mitigated but not eliminated: the three-model panel, the researcher hand-scoring layer, and the five-rater validation provide three distinct validity claims, yet consensus agreement against hand-scoring cleared at only κ = 0.704, and within-family scorer divergence on engagement redirection remained (0.42 and 0.40 against the Gemma comparisons).
Connected Concepts
- AI Sycophancy
- Conversational AI
- Large Language Models (LLMs)
- Generative AI
- Trust
- Trust Calibration
- Human-in-the-Loop
- Human AI Collaboration
- AI Literacy
- Technology Adoption Models
- Equity
- Digital Divide
- Anxiety and Stress
- Bias Mitigation
- Ethics
- Well-Being
- Educational AI Policy
- Research Methods in AIED
Connected Articles
- The Hidden Cost of Contextual Sycophancy: an AI Literacy Intervention in Human-AI Collaboration — Sycophancy in AI output and its implications for AI literacy
- Unpacking ethics-domain of intelligent-TPACK scale in relation to in-service teachers' trust and distrust — Teacher trust and distrust of AI as a dimension of ethical practice
- The Role of Artificial Intelligence Anxiety and Attitudes Toward Artificial Intelligence in University Students' Job Finding Anxiety — AI anxiety and employment/occupational concern among learners
Citation
Not for People Like Me: How Frontier AI Models Redirect Skeptical Rural School Staff — Rossmiller, Z. (2026). Computers and Education: Artificial Intelligence, 11, 100659.