$ open ~/projects/trust-forge
Trust Forge
Agent benchmark & training platform
Central evaluation engine for agent benchmarks, tests, and LoRA fine-tuning — Godot operator console, FastAPI control plane, versioned task suites, and Docker SWE-Bench sandboxes. Built to benchmark Ember over A2A, run Ollama/OpenAI evals directly, and fine-tune small models for on-device NPC workflows.
- Godot console — 11 screens for runs, training, reports, and NPC Lab
- FastAPI backend — WebSocket control plane, job queue, leaderboard
- Benchmark suites — Ember A2A, MINT, SWE-Bench, BFCL, NPC persona evals
- LoRA pipeline — trajectory export, adapter training, PEFT re-benchmark
- CLI + HTTP API — catalog, runs, GGUF export, persona chat
- Docker SWE sandbox — pytest-based code repair evaluation