How well frontier models complete their assigned tasks while a trusted monitor watches for sabotage. Usefulness counts honest tasks completed; safety counts attacks stopped at a fixed auditing budget.