more tests passed = better model.
But if tests or feedback are accessible, models may learn to game them.
We introduce CapCode to detect suspiciously high scores, and CapReward to discourage them during RL.
🧵1/10
more tests passed = better model.
But if tests or feedback are accessible, models may learn to game them.
We introduce CapCode to detect suspiciously high scores, and CapReward to discourage them during RL.
🧵1/10
With the manufacturer datasheet, a connection could be made between any suitable pin...
2/
With the manufacturer datasheet, a connection could be made between any suitable pin...
2/
Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests
By Lodkaew, Ackermann, Nishimori et al
Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests
By Lodkaew, Ackermann, Nishimori et al