Nine Bugs in My Own Evaluation Harness. Every One Made My Results Look Better.
I built an evaluation harness. I found nine bugs in it. Every single one would have made my results look better than they were. Nine out of nine, all pointing the same way. That is not a coincidence, and I do not think it is specific to me or to my project. I think it is a structural problem with any evaluation you build for yourself. Here is the mechanism, before any of the evidence. Debugging is triggered by surprise You do not audit numbers. You audit numbers that bother you. When a result comes back disappointing, you go looking for the reason. You check the setup, you re-run it, you add logging, you find the bug. The bug gets fixed and the number moves. When a result comes back good, none of that fires. Nothing feels wrong. There is no surprise to investigate. You write it up. So the filter that removes measurement bugs from your work is applied unevenly. Hard against results you dislike. Softly against results you like. Every pass through that filter removes more unflattering bugs than flattering ones. Run that loop for a few days and your instrument has drifted in one direction, and nothing in your process is designed to notice. The bugs that survive to publication are disproportionately the ones that helped you. Not because anyone was dishonest. Because they never triggered the thing that catches bugs. I knew this argument in the abstract before I started. It did not stop me from writing nine of them. What the harness did Only enough context for the bugs to make sense. Mutation testing changes your source code in small ways. Flip a comparison. Change a constant. Delete a raise . Then it runs your test suite and checks whether anything failed. def withdraw(balance, amount): - if amount <= 0: + if amount < 0: raise ValueError("amount must be positive") If the suite stays green, that is a fault your tests cannot detect. Coverage tells you a line ran. This tells you whether anything would have complained if the line were wrong. Those are very different questions