Skip to content
View mottopanikeiku's full-sized avatar

Highlights

  • Pro

Block or report mottopanikeiku

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
mottopanikeiku/README.md

I’m building Tenth Man Labs. We put probabilities on geopolitical events, with the sources behind each number and a dated record of how it moved.

My background is GPU systems. Lately most of my work is evaluation and reinforcement learning.

Start here

  • attention-numerics: when rotating queries and keys before FP8 rounding makes attention worse. With the real FlashAttention-3 FP8 kernel on an H100, rotation cuts Qwen2.5-7B’s HellaSwag accuracy from 80.9% to 45.5%; centering the keys first removes every measured loss of two points or more across eight models. Interactive explanation.
  • eval-power: how many benchmark items it takes to tell two LLMs apart, and how often a small pilot gets that number wrong. In a live test on six small models, plans committed before collecting fresh answers detected 23 of 27 differences. Calculator.
  • seed-power: the same question for RL seeds. On more than 77,000 fresh training runs, studies planned from small pilots for 80% power detected the difference 56% of the time with linear policies and 45% with neural PPO.
  • cpu-decode: a from-scratch C++ int8 decoder for Qwen2.5-0.5B. At short context it reaches 83% of the laptop's memory-read ceiling and, at its best thread count, beats llama.cpp; at long context llama.cpp is clearly faster.
  • control-clock: seconds from a fresh Python process to a policy that passes CartPole or Acrobot. Random search over linear policies beat every PPO on CartPole; tuned PPO won Acrobot; a cloud GPU was slower once startup counted.
  • alignmenttax: does instruction tuning trade calibration for truthfulness? In seven base/instruct pairs, calibration got worse in six, but only four gained accuracy in exchange.
  • quantile-cycles: a counterexample in distributional RL, checked in exact rational arithmetic and in Lean.
  • verge-lab: picking preference pairs from multi-aspect judge scores, and abstaining when the aspects disagree. Neither agreement with human preferences nor DPO training on the selected pairs beat a simple score-gap rule.
  • branchpilot: a learned stopping rule for self-consistency sampling. Simple agreement rules won on GSM8K and on a MATH-500 holdout.
  • faultline: recurrent PPO agents that have to run a cheap test before an expensive repair. Across 375 seeds per curriculum, training only on ambiguous faults beat random sampling by 10 points but lost to a difficulty curriculum by 6.
  • heliostune: Triton autotuning across four NVIDIA GPUs. Tuning data from other GPUs didn’t help; on the H100 the limit was the kernel set, and adding split-K and persistent kernels won 9 of 96 workloads against torch.matmul, up from none.

Merged fixes in PyTorch, Ray, Sentence Transformers, SGLang, and Celery. More at mottopanikeiku.github.io.

Pinned Loading

  1. EzPC EzPC Public

    Forked from mpc-msri/EzPC

    C++

  2. py-closewat py-closewat Public

    Python port of the closewat C tool for water molecules in X-ray PDB structures, with C-vs-Python reference tests

    Python

  3. phantom-fhe phantom-fhe Public

    Forked from encryptorion-lab/phantom-fhe

    PhantomFHE: A CUDA-Accelerated Homomorphic Encryption Library

    Cuda