Skip to content

Size Brainstore cache as a percentage of the cache volume - #108

Draft
Jeff McCollum (jeffmccollum) wants to merge 2 commits into
mainfrom
writer-cache-75pct
Draft

Jeff McCollum (jeffmccollum) wants to merge 2 commits into
mainfrom
writer-cache-75pct

Conversation

@jeffmccollum

Copy link
Copy Markdown
Contributor

Summary

Helm version of braintrustdata/terraform-aws-braintrust-data-plane#420. Brainstore's object store cache is now sized as a percentage of the cache volume, like the AWS Terraform module, instead of a fixed 1000Gi default. Readers and fast readers use 90%. Writers use 75%.

  • New brainstore.<role>.objectStoreCacheFileSizePercent (90 / 90 / 75). objectStoreCacheFileSize now defaults to "".
  • When objectStoreCacheFileSize is empty, the container starts through /bin/sh -c. It reads the size of the filesystem at cacheDir and caps it at the smallest of volume.sizeLimit, volume.size, ephemeralStorage.request and ephemeralStorage.limit that is set. It then exports BRAINSTORE_OBJECT_STORE_CACHE_FILE_SIZE (whole Gi, rounded down) and execs brainstore web. The size it picks is logged at startup. Under 1Gi it exits with an error.
  • When objectStoreCacheFileSize is set, nothing changes: same brainstore web command, same ConfigMap key.
  • The percentage has a default in the template too, so helm upgrade --reuse-values from older chart versions still works.
  • GKE examples drop their fixed sizes. google-standard had a 1000Gi cache on a 200Gi volume.
  • README has a new "Brainstore Cache Sizing" section with per-platform behavior: GKE Autopilot and Standard, EKS (including Auto Mode), and AKS with or without Azure Container Storage.

Upgrade impact

  • Deployments that relied on the 1000Gi default get a computed size, and their Brainstore pods restart, because the pod template changes.
  • Deployments that set objectStoreCacheFileSize, or that upgrade with --reuse-values, see no change.
  • When the cache is an emptyDir on a shared node disk and no size is set, the percentage applies to the whole node disk. Brainstore pods need dedicated nodes, or one of the sizes above set. This is documented in the README.
  • No Chart.yaml version bump, per AGENTS.md.

Validation

  • ./test.sh passes: 362 unit tests, multi-cloud rendering, and helm lint. The new brainstore-cache-file-size_test.yaml covers fixed vs percentage sizing, per-role defaults, every cap source, EKS ephemeralStorage, Azure Container Storage, and bad input. 13 of its first 15 tests fail against main.
  • Ran the rendered startup script in the real public.ecr.aws/braintrust/brainstore:v2.17.0 filesystem (chroot, uid 1000, 10GiB tmpfs at cacheDir, a stand-in brainstore binary):
    • writer: 7Gi
    • reader: 8Gi
    • writer with sizeLimit: 4Gi: 2Gi
    • volume.size: 100Gi larger than the disk: the disk size is used
    • sizeLimit: 1Gi: exits 1 with an error
  • Not yet deployed to a real GKE, EKS, EKS Auto Mode or AKS cluster.

🤖 Generated with Claude Code

Mirrors terraform-aws-braintrust-data-plane#420. When
objectStoreCacheFileSize is unset (the new default), each Brainstore pod sets
BRAINSTORE_OBJECT_STORE_CACHE_FILE_SIZE at startup to
objectStoreCacheFileSizePercent of its cache volume: 90% for readers and fast
readers, 75% for writers. The volume size is the filesystem at cacheDir,
capped by volume.sizeLimit and volume.size when set.

An explicit objectStoreCacheFileSize keeps the existing behavior: Brainstore
starts directly and reads the size from its ConfigMap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On EKS the documented storage budget is ephemeralStorage.request, so the
computed cache size now also respects ephemeralStorage.request and
ephemeralStorage.limit. Document per-platform behavior in the README.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant