Execute and monitor LLM testing
No test runs yet
This prompt will be sent to all test models to guide their responses
No challenge sets configured. Add challenge sets first.
No models configured. Add models first.
Delay between each API call to avoid rate limits (0-1000ms)
This model will evaluate responses and provide scores (required)
Instructions for the moderator (e.g., evaluation criteria, language preference)
Minimum moderator score to consider a response as "passed" (0-100)