Skip to content

epic(bench): standardize BenchExec and Divan measurement and retire custom benchmark harnesses #959

Description

@DecisionNerd

Objective

Standardize performance measurement on BenchExec for command/process-tree workloads and Divan for in-process benchmarks. Retire duplicated benchmark machinery with bounded semantic-parity evidence, preserving correctness tests and historical results. This small epic is executable and closable independently of #900; #952 remains its parent.

Maintainer decisions (2026-09-19)

Scope and ownership

  1. Policy, inventory, testing docs and CI enforcement (test(bench): define and enforce the BenchExec and Divan measurement policy #1484): a native child owns this repository-wide scope, additional to the original ladder-retirement issue. Inventory producers and consumers, measurement boundaries and dispositions: benchmark authority, diagnostic attribution, deadline/control logic or unused code.
  2. Traversal/MERGE migration to Divan (refactor(bench): migrate traversal and MERGE scaling measurements to Divan #1485): a native child owns these in-process workloads, also outside the original scope. Preserve mutable fixture state, warmed traversal intent, timed setup/teardown, graph/hop parameters and deterministic assertions.
  3. Lifecycle/GDC orchestration and bounded parity: retained here. Inspect certification/GDC runners and BenchExec adapters/consumers. Remove unused timers and duplicate benchmark authority; preserve or consolidate scoped shared diagnostics. Retain load/query/validation semantics, failures, cleanup and evidence on bounded tiny/shadow fixtures and policy tests. Do not split an in-memory GDC workload across processes merely to get separate timers. Establish supported framework measurement boundaries before replacing diagnostic-derived thresholds. Reconcile cgroup memory versus process RSS names where required by the migration; preserve schema compatibility. No live GDC suite execution and no test(scale): complete the S18-S26 lifecycle ladder on OVHC-AGENCY #900 ladder run are closure prerequisites.
  4. Existing owners remain canonical: perf(bench): make the ingest gates two-sided so improvements ratchet automatically (#1387 D1/D5) #1476 owns ingest gate/ratchet implementation and migration of its custom wall/CPU inputs; test(tck): the performance baseline is from a CI runner, so the perf gate warns on every local run #1467 owns TCK comparisons and measurement migration. Reuse their work instead of duplicating it. This epic verifies only their policy-conformance outcomes, not unrelated optimization or outlier repairs.

Keep the inventory finite and concrete: traversal/MERGE, Divan ingest including its alternate gate mode, certification/lifecycle, GDC and TCK, plus independently discovered performance-gate consumers in the same audit. Additional issues require a verified blocker or independently reviewable XL concern, not one issue per timer.

Execution order

The policy/inventory child #1484 is ready now, subject to the repository WIP limit. The Divan migration #1485 is natively blocked by #1484 and follows that policy contract. Bounded lifecycle/GDC diagnosis and fixture preparation can proceed immediately; perform its implementation against the agreed policy. Merge focused changes promptly and refresh the next candidate against current main.

Native dependencies own execution state. #956/#957/#1112 are delivered foundations, not work to repeat. #1112 already separated structural retirement, prefix parity and full completion; adapt its consumer contract to this maintainer decision without claiming an incomplete ladder is complete.

Acceptance criteria

Completion scenarios

  1. Given a new custom benchmark sampler or diagnostic-based performance gate, when policy and evidence tests run, then they fail with an actionable location; deadlines and approved shared diagnostics remain usable.
  2. Given equivalent fresh/warmed fixtures, when Divan workloads execute, then the same work and deterministic assertions hold and timing comes from Divan.
  3. Given legacy and new bounded lifecycle fixtures, when normalized evidence is compared, then behavior agrees or every difference is explained, and failure/cleanup cases remain fail-closed. Live GDC suite execution is not required.
  4. Given all bounded migration outcomes are merged and green but test(scale): complete the S18-S26 lifecycle ladder on OVHC-AGENCY #900 is incomplete and no GDC live lane has run, when epic(bench): standardize BenchExec and Divan measurement and retire custom benchmark harnesses #959 readiness is evaluated, then this epic can close while test(scale): complete the S18-S26 lifecycle ladder on OVHC-AGENCY #900 and epic(bench): externalize scale certification and add GDC benchmark suites #952 continue to require their own real scale evidence.

Relationships and non-goals

Parent #952 retains the original completed-ladder comparison obligation using the single #900 OVHC-AGENCY run, including all semantic dimensions above. Neither that obligation nor #900 blocks this epic. Keep delivered #956/#957/#1112 relationships for provenance; Fly #958 is optional historical tooling.

No new framework, product-only-for-benchmarks API, numerical optimization target, weakened durability/correctness, new release cascade or scale execution. OVHC-AGENCY remains the scale acceptance host. Closing this epic proves standardized measurement and bounded migration parity, not billion-edge throughput or complete scale certification.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    testingTest coverage and testing infrastructure

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions