Back to projects

lvl

done

A local-first arena for evaluating model-and-harness combinations through chess matches and puzzle batches.

type
agent evaluation arena
for
AI labs

what it does

  • Paired matches and custom task packs
  • Browser traces, replay, and PGN output
  • Latency and cost tracking
  • Stockfish scoring and local persistence
technical details
  • TypeScript
  • React
  • Vite
  • Express
  • Playwright
  • SQLite
  • Stockfish
  • OpenRouter