Test Runs

Execute and monitor LLM testing

Previous Test Runs

No test runs yet

Create New Test Run

This prompt will be sent to all test models to guide their responses

No challenge sets configured. Add challenge sets first.

No models configured. Add models first.

50 ms

Delay between each API call to avoid rate limits (0-1000ms)

This model will evaluate responses and provide scores (required)

Instructions for the moderator (e.g., evaluation criteria, language preference)

70 points

Minimum moderator score to consider a response as "passed" (0-100)