The best coding agent clears less than half of a new benchmark built to test whether language models can build static-analysis checkers from scratch.