Skip to main content
Run multiple configured models concurrently:
Each contestant is bound to its declared provider/model. The judge receives only the task and contestant outputs and scores correctness, completeness, testability, and safety. Failed contestants are retained in the report, and a judge failure is explicit rather than silently selecting an answer. Natural language works: “race several agents on a coding task.” Use the result as a recommendation; Arka does not write, deploy, register, or submit anything automatically from a race winner.