EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models

arXiv نُشر في تم التحديث الذكاء الاصطناعي وتعلّم الآلة
سجّل الدخول للحفظ

الأصول والمواضيع المتأثرة

يُعرض هذا المقال بلغته الإنجليزية الأصلية.

لماذا يهم

التحليل معروض بالإنجليزية · الترجمة العربية قيد الإعداد

A new benchmark, EvalDetectBench, was introduced to measure 'evaluation awareness' in frontier large language models (LLMs), which refers to models recognizing when they are being evaluated. The benchmark highlights methodological biases in existing evaluations, such as the influence of transcript identity and elicitation prompts, which can skew model rankings and measurement validity.

  • introduction of EvalDetectBench as a new benchmark for measuring evaluation awareness in LLMs
  • identification of methodological biases (11.25% variance from transcript identity and elicitation prompt performance disparities) that undermine evaluation validity
  • benchmark's use of Inspect-compatible evaluations and curated transcript suite for current frontier system-card evaluations

التأثير المتوقع على السوق

محايد الثقة 85% كيف تُقرأ نسبة الثقة الأفق الزمني: المدى المتوسط الأثر: الأعلى

The benchmark may affect public AI companies exposed to LLM evaluation frameworks by increasing scrutiny of their model performance metrics, potentially influencing investor confidence in their safety and reliability claims. Specifically, companies reliant on LLM benchmarks for credibility (e.g., those with frontier models) could face reputational or operational risks if their models exhibit high evaluation awareness.

المخاطر

  • article does not provide evidence of immediate market reactions or capital flows tied to the benchmark
  • no direct financial metrics or earnings implications are discussed, limiting near-term market impact assessment

مسار الأدلة

الأدلة
المصدر arXiv
الادّعاء EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models
الأصول المتأثرة GOOGL, MSFT, NVDA, AMZN, META
استنتاج الذكاء الاصطناعي محايد · 85%
أُنشئ في 2026-09-03 04:00

مصدر التحليل بالذكاء الاصطناعي

حُلِّل بواسطة Mistral Small Latest المنهجية v1.0 أُنشئ في
المعرّفات التقنية
وسم المزوّد
mistral-small-latest
إصدار التحليل
mistral-small-latest
معرّف المقال
126647
الإطار الزمني
24h

دورة حياة التوقّع

  • Mistral Small Latest GOOGL محايد 85% 24h
    أُنشئ في 6 س 24 س مُتحقق منه
  • Mistral Small Latest MSFT محايد 85% 24h
    أُنشئ في 6 س 24 س مُتحقق منه
  • Mistral Small Latest NVDA محايد 85% 24h
    أُنشئ في 6 س 24 س مُتحقق منه
  • Mistral Small Latest AMZN محايد 85% 24h
    أُنشئ في 6 س 24 س مُتحقق منه
  • Mistral Small Latest META محايد 85% 24h
    أُنشئ في 6 س 24 س مُتحقق منه

يُسجَّل وقت النشر، ويُقيَّم تلقائياً بمجرد انتهاء النافذة الزمنية — دون أي تعديل.

المصدر الأصلي

arXiv:2609.01611v1 Announce Type: new Abstract: Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness. If models behave differently in evaluations than in deployment, this undermines the validity of evaluation results, which are a crucial component of current AI safety frameworks. We introduce EvalDetectBench, an open pipeline and benchmark for measuring evaluation awareness that works with any Inspect-compatible evaluation, allowing practitioners to test against current and future benchmarks. EvalDetectBench ships with a newly curated transcript suite covering current frontier system-card evaluations and diverse deployment sources. The benchmark serves two purposes: measuring how reliably frontier LLMs recognize that they are being evaluated, and assessing how detectable individual benchmarks are as evaluations. We identify two methodological choices in the existing literature that introduce systematic bias: the identity of the model that generated the deployment transcripts accounts for 11.25% of measurement variance and can reorder model rankings; and elicitation prompts selected for high performance on one model can perform near chance on others. EvalDetectBench corrects for both via per-model probe calibration and a stratified generator-harmonisation procedure.

اقرأ المقال كاملاً على arXiv

المقال الأصلي منشور بواسطة arXiv في 3 سبتمبر 2026. التحليل والرؤى المقدمة من AnalystMarkets AI.

المزيد من سردية GOOGL

أداء هذا النموذج على أخبار مشابهة

Mistral Small Latest · 36.7% صحيحة عبر 1083 توقّعاً مُقيَّماً على الأسهم اطّلع على السجل الكامل