You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Build independently labeled, versioned cases from authorized traces and actual outcomes. Include legitimate test edits, unsafe quotations/refusals, denied attempts, missing telemetry, ordinary delegation and unusual but permitted behavior; keep source-unit independence and development/final separation.
Attack the monitor's own evidence boundary: injected instructions, misleading framing, omitted/forged evidence, long distractors and cross-run combinations. Preserve actual policy, target, timestamp, actor and attempted/denied/completed outcomes; never substitute model-generated explanations for independent truth.
Compare existing deterministic checks, deep analysts, bounded native classifiers and escalation combinations. Report per-category false-positive/false-negative rates, unresolved coverage, incident prevalence, total cost/latency and matched baseline; freeze each policy before final assessment.
Prove authorized storage and actual Runtime/Platform hook placement before any operational adoption. Keep enforcement outside the evaluated agent, retain required deterministic controls and test failure/revocation/resume; no automatic permissions or production promotion from a classifier score.
Acceptance: reproducible held-out behavioral results and a separate adaptive-adversary report attached to exact model/question/policy versions, with uncertainty and unsupported claims explicit. Existing auditEvaluator/auditProbabilityPolicy, reward-hacking diagnostics, control-integrity checks and AnalystRegistry remain the owners; no new ledger or detector runtime.
The implementation recipe and mapping tests do not close this issue. Jev 1.13's own documentation warns adversarial state can steer answers (https://docs.typesafe.ai/model-jaggedness/jev-1.13). Native confidence describes a distribution statistic, not calibrated malicious intent. Unknown is not clean, anomaly is not malice, and correlated question probabilities must not be multiplied into a run-level guarantee. No live classifier quality or deployed safety effectiveness has been demonstrated in this pass.
Parent #768; extends #771 calibration and #772 trace-analysis adoption. User-requested application: reward gaming, evaluator tampering, safety-policy and authorization violations, data-boundary breaches, injection-following, contradicted completion claims, oversight bypass, matched-baseline anomalies and cross-run behavior.
Acceptance: reproducible held-out behavioral results and a separate adaptive-adversary report attached to exact model/question/policy versions, with uncertainty and unsupported claims explicit. Existing auditEvaluator/auditProbabilityPolicy, reward-hacking diagnostics, control-integrity checks and AnalystRegistry remain the owners; no new ledger or detector runtime.
The implementation recipe and mapping tests do not close this issue. Jev 1.13's own documentation warns adversarial state can steer answers (https://docs.typesafe.ai/model-jaggedness/jev-1.13). Native confidence describes a distribution statistic, not calibrated malicious intent. Unknown is not clean, anomaly is not malice, and correlated question probabilities must not be multiplied into a run-level guarantee. No live classifier quality or deployed safety effectiveness has been demonstrated in this pass.