Compare
The front door. Run any model from any source live on a catalog dataset or your own prompts, on text or image, then settle the unscored tasks with a blind preference bracket.
- ✓Curated, OpenRouter, frontier, and self-hosted vLLM in one list
- ✓Quality only where references exist; latency, TTFT, throughput always measured
- ✓One blind single-elimination bracket per prompt, votes become routing labels
Evolve
GEPA search over instruction, demos, model, and policies under an explicit quality floor. Catalog datasets, or your own goal and rubric scored by an LLM judge or checklist QWK.
- ✓Baseline vs evolved report with a Pareto chart you can regenerate offline
- ✓Train and val are fair game; test is reported once and refused if reused
- ✓Tokenizer fertility caps demos: watch demos_requested vs demos_fitted
Deploy
Serve self-hosted models on your own GPU host over SSH: measure, start, health, benchmark, resumable terminal. Then put the served endpoint back in Compare against the frontier.
- ✓Self-hosted relative cost comes from measured throughput, not published pricing
- ✓FP8 vs bf16 always appears in result caveats
- ✓GPU host config stays in a gitignored root .env