Back to platforms

AI Architect Practice Arena

LLM-as-judge mock interviews — 35 playbook questions, dual grading.

Sectioned system-design forms, STAR behavioral practice, and reasoning trade-offs — graded by OpenAI and Anthropic against Staff+/Principal rubrics. Bring your own API key; keys never touch our servers.

In one sentence

Mock system design and behavioral interviews with dual LLM judges.

Decision

Per-question rubrics over generic chat — grading matches playbook depth.

Measured signal

139/140 live cases passed · judge disagreement shown, not averaged away.

Honest limitation

Requires user's OpenAI or Anthropic API key; no server-side key storage.

  • 35/35 interview playbook coverage across three rubric formats
  • Dual-judge grading with disagreement surfaced as trust signal
  • 139/140 live calibration cases passed across two independent runs
  • BYOK architecture — zero org-side API cost at any scale
Multi-agent reference topologyArchitecture
EXPERIENCECONTROL PLANEMODEL PLANEOBSERVABILITYUser / OpsPolicy & GuardrailsAegisAIOrchestratorLangGraphSpecialist AgentsHybrid RAGTools / APIsEvaluation GatesGateway / HITLLLM Gatewayaegis-llm-gatewaySemantic Cacheaegis-semantic-cacheTraces · Audit · FinOpstrace-linked LLMOps

Next.js · Vercel · OpenAI · Anthropic APIs