Skip to content

[Typed evaluation P1] Validate behavioral risk monitors against independent and adversarial traces #774

Description

@drewstone

Parent #768; extends #771 calibration and #772 trace-analysis adoption. User-requested application: reward gaming, evaluator tampering, safety-policy and authorization violations, data-boundary breaches, injection-following, contradicted completion claims, oversight bypass, matched-baseline anomalies and cross-run behavior.

  • Build independently labeled, versioned cases from authorized traces and actual outcomes. Include legitimate test edits, unsafe quotations/refusals, denied attempts, missing telemetry, ordinary delegation and unusual but permitted behavior; keep source-unit independence and development/final separation.
  • Attack the monitor's own evidence boundary: injected instructions, misleading framing, omitted/forged evidence, long distractors and cross-run combinations. Preserve actual policy, target, timestamp, actor and attempted/denied/completed outcomes; never substitute model-generated explanations for independent truth.
  • Compare existing deterministic checks, deep analysts, bounded native classifiers and escalation combinations. Report per-category false-positive/false-negative rates, unresolved coverage, incident prevalence, total cost/latency and matched baseline; freeze each policy before final assessment.
  • Prove authorized storage and actual Runtime/Platform hook placement before any operational adoption. Keep enforcement outside the evaluated agent, retain required deterministic controls and test failure/revocation/resume; no automatic permissions or production promotion from a classifier score.

Acceptance: reproducible held-out behavioral results and a separate adaptive-adversary report attached to exact model/question/policy versions, with uncertainty and unsupported claims explicit. Existing auditEvaluator/auditProbabilityPolicy, reward-hacking diagnostics, control-integrity checks and AnalystRegistry remain the owners; no new ledger or detector runtime.

The implementation recipe and mapping tests do not close this issue. Jev 1.13's own documentation warns adversarial state can steer answers (https://docs.typesafe.ai/model-jaggedness/jev-1.13). Native confidence describes a distribution statistic, not calibrated malicious intent. Unknown is not clean, anomaly is not malice, and correlated question probabilities must not be multiplied into a run-level guarantee. No live classifier quality or deployed safety effectiveness has been demonstrated in this pass.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions