Back to platforms

ModelForge

Model Plane flagship — which weights, where, with what proof.

Hire-facing Model Plane for SLM bake-offs, PEFT receipts, CUDA vLLM metrics, and LLM gateway enforce+record. Composes DomainForge + upstream vLLM + aegis-llm-gateway (ADR-034).

In one sentence

Model Plane control surface — SLM, PEFT, CUDA vLLM, LLMOps — peer to the agent spine.

Decision

One ModelForge flagship beats burying PEFT/vLLM in a teaching drawer (ADR-034).

Measured signal

Live https://modelforge-gamma.vercel.app/api/v1/posture — PEFT + CUDA vLLM + SLM + gateway all ready; peft_gpu.json / vllm_cuda.json from GCP Tesla T4 (cuda=true).

Honest limitation

PEFT receipt is a T4 fp16 LoRA micro-run (honest), not a DomainForge 7B S0/S3/S4 hire-depth ladder. vLLM metrics are TinyLlama on T4, not always-on production serve.

  • Honest /api/v1/posture (ready vs smoke vs planned)
  • Receipt gallery for PEFT · CUDA vLLM · SLM bake-off
  • Composes DomainForge train + gateway route + vLLM serve path
  • Panel scripts for buy vs RAG vs PEFT vs self-host

Read related ADR →

Multi-agent reference topologyArchitecture
EXPERIENCECONTROL PLANEMODEL PLANEOBSERVABILITYUser / OpsPolicy & GuardrailsAegisAIOrchestratorLangGraphSpecialist AgentsHybrid RAGTools / APIsEvaluation GatesGateway / HITLLLM Gatewayaegis-llm-gatewaySemantic Cacheaegis-semantic-cacheTraces · Audit · FinOpstrace-linked LLMOps

Next.js · FastAPI · TRL/PEFT (via DomainForge) · vLLM · Vercel