The best coding agent clears less than half of a new benchmark built to test whether language models can build static-analysis checkers from scratch.
Some results have been hidden because they may be inaccessible to you
Show inaccessible resultsSome results have been hidden because they may be inaccessible to you
Show inaccessible results