Research Article
Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration
Synthesis: EchoPrompt introduces a training-free zero-shot detector for LLM-generated text that exploits the latent prompt dependency inherent in machine-generated content. By restoring a generic assistant-response prefix and measuring likelihood gain differences between instruction-tuned and base models, EchoPrompt achieves state-of-the-art detection performance without training. This approach has direct implications for academic integrity in educational contexts, where LLM-generated text detection is increasingly critical for maintaining assessment validity.
Detection Mechanism
EchoPrompt is built on the insight that machine-generated text is conditioned on an upstream prompt, and this hidden dependency can be partially reactivated. The detector:
- Prepends a unified generic prefix that mimics an assistant-response context
- Measures induced likelihood gain with an instruction-tuned model
- Calibrates against the corresponding base model to control for model-specific biases
- Aggregates likelihood differences into a score quantifying latent prompt dependency
This training-free approach contrasts with existing zero-shot detectors that rely purely on probability-based statistical discrepancies without modeling the generation mechanism.
Key Findings
- State-of-the-art zero-shot detection: EchoPrompt outperforms existing zero-shot detectors across multiple evaluation settings
- Robustness: Strong performance maintained across challenging scenarios including domain shift and paraphrasing attacks
- No training required: The detector is fully training-free, relying only on access to instruction-tuned and base model pairs
- Educational relevance: Directly addresses growing concerns about educational misuse of LLMs for generating assignments, essays, and exam responses
Latent Prompt Dependency and the EchoPrompt Score
EchoPrompt rests on a simple empirical observation: because modern large language models are post-trained under a global instruction, the text they produce implicitly "remembers" that it was written as a response. Even when the original prompt is stripped away, machine-generated passages align more naturally with a restored assistant-style context than human-written text does. To operationalize this, EchoPrompt prepends a task-agnostic prefix — "You are a helpful, versatile, and intelligent AI assistant…" — to approximate the generic condition under which AI output is produced.
A direct likelihood measure under the instruction-tuned model is not discriminative on its own, since high-frequency tokens, common phrases, and raw fluency inflate token probabilities for both human and machine text. EchoPrompt therefore calibrates the instruction-tuned model's estimate of the restored sequence [cg; X] against the corresponding base model's estimate of the original text X. The base model captures only the marginal linguistic regularities of open-domain text, so the difference suppresses shared fluency effects and isolates the extra advantage a passage receives under assistant-style conditioning. Averaging this token-level gap yields a stable sequence-level score, which is compared against a threshold to classify the passage.
Robustness and Evaluation
EchoPrompt was evaluated on three public detection benchmarks — DetectRL, RealDet, and RAID — spanning multi-domain, multi-LLM, and multi-attack splits. Paired base/instruct proxies from Qwen2.5, Llama-3.2, Llama-3.1, Llama-3, and Falcon families were used, with AUROC and F1 as the primary metrics. Under the Llama-3-8B proxy, EchoPrompt ranked first on both metrics across all three benchmarks, improving over the strongest training-free baseline (IRM) by 0.69% AUROC and 2.64% F1 on average, and by far larger margins over the best training-based detector.
The method proved robust to adversarial transformation, obtaining the best scores in four of five attack groups and improving over IRM by 1.37% F1 on average under direct prompting, perturbation, prompt attacks, and data mixing. Its strongest performance came precisely where likelihood-, entropy-, and rank-based baselines falter, because it detects generation-style dependency rather than isolated token statistics. An ablation study confirmed that the restored prompt–response framing (context clause A) drives most of the gain — the full prefix improved AUROC by up to 14.73% over the empty-prompt setting — and that performance remains strong across proxy families and scales, with inference latency under 0.26 seconds per sample.
What this means for practice
- Designers. Deploy a training-free detector instead of retraining one: EchoPrompt needs only an instruction-tuned and base model pair and classifies a sample in under 0.26 seconds, so screening can follow commercial LLM releases without a new labeled corpus.
- Designers. Treat the score as one signal, never as proof of authorship: the authors flag false positives that wrongly flag human writing and false negatives that miss generated text as the harms that matter most in high-stakes settings.
- Administrators. Pair detection with human oversight and transparent AI Governance rather than policy that leans on the score alone; robustness to paraphrase attacks comes from modeling the prompt-response relation, which is not evidence about who wrote a passage.
- Designers. Budget proxy and prefix choices as maintenance work: detection depends on the choice of base/instruction-tuned proxy family, and the generic prefix is empirically tuned rather than proven optimal.
Limitations
- Proxy-dependent detection. Like other zero-shot detectors, EchoPrompt depends on the choice of proxy family, and its current prefix is empirically tuned rather than proven globally optimal.
- False positives and false negatives. Automated detection may wrongly flag human writing as machine-generated or miss generated content — harms that are especially consequential in high-stakes settings.
- Not evidence of authorship. EchoPrompt is best treated as an auxiliary signal rather than definitive evidence of who wrote a passage.
Citation
Bao, H., Ren, Y., Cao, Y., You, J., Fang, F., & Wang, S. (2026). Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration. v1.