Model Plane
ModelForge
Model Plane flagship — which weights, where, with what proof.
Hire-facing Model Plane for SLM bake-offs, PEFT receipts, CUDA vLLM metrics, and LLM gateway enforce+record. Composes DomainForge + upstream vLLM + aegis-llm-gateway (ADR-034).
In one sentence
Model Plane control surface — SLM, PEFT, CUDA vLLM, LLMOps — peer to the agent spine.
Decision
One ModelForge flagship beats burying PEFT/vLLM in a teaching drawer (ADR-034).
Measured signal
Live https://modelforge-gamma.vercel.app/api/v1/posture — PEFT + CUDA vLLM + SLM + gateway all ready; both peft_gpu.json and vllm_cuda.json are real Mistral-7B-Instruct-v0.3 receipts on a rented L4 — real QLoRA SFT + DPO training (peft_gpu.json) and real upstream vLLM serving (vllm_cuda.json, 13.74 tok/s, TTFT p50 371.67ms).
Honest limitation
PEFT receipt reports real training config/timing, not a quality/win-rate score — DomainForge's S0-S4 eval harness isn't wired to real adapter inference yet (see the receipt's own known_gaps). vLLM metrics are a single dated benchmark run, not an always-on production serve claim.
- Honest /api/v1/posture (ready vs smoke vs planned)
- Receipt gallery for PEFT · CUDA vLLM · SLM bake-off
- Taxonomy glassbox — LoRA · QLoRA · Multi-LoRA · classical ML lane
- Composes DomainForge train + gateway route + vLLM serve path
- Panel scripts for buy vs RAG vs PEFT vs self-host
Video walkthrough — 10 min (planned)
ModelForge: Model Plane posture and CUDA receipts
Walk /api/v1/posture, SLM bake-off table, peft_gpu.json and vllm_cuda.json — buy vs RAG vs PEFT vs self-host without overclaiming.
Recording in progress. Subscribe on YouTube to get notified when this walkthrough publishes.
Subscribe @venkat-aiArchitecture diagram
Next.js · FastAPI · TRL/PEFT (via DomainForge) · vLLM · Vercel