Charlie Fuller
I build AI that organizations trust
Not demos. Production platforms that evaluate AI agents, govern legal workflows, and automate operations — deployed to real users, backed by 1,050+ tests and 132 visual explainers.
AI Evaluation
AESOP Transformation OS
The core problem in enterprise AI isn't building agents — it's knowing whether they work. AESOP is a full-lifecycle platform that discovers, builds, evaluates, and monitors AI agents. Its evaluation engine uses 20+ independent scorers because a model can't grade its own homework. Deployed across multiple organizations, ported to legal, used to independently evaluate third-party AI agents in regulated environments.
See the platform →
Legal AI
Legal AI Operating System
A governed platform for legal AI adoption. Five-layer architecture: knowledge foundation, governance framework, AI functions, program operations, organizational model. The breakthrough is independent agent monitoring — external evaluators score AI outputs across multiple dimensions. The AI can't grade itself. This platform does.
See the platform →
Voice & Workflow AI
Conversational AI That Configures Itself
A five-minute voice interview replaces the onboarding form. The user describes their role, priorities, and communication style in their own words — and the system configures itself around them. Behind the scenes: structured question flows with branching logic, real-time transcription, and an agent orchestration layer that translates natural conversation into a configured AI assistant. Voice pipeline is fully serverless (ElevenLabs + Railway + Supabase). Production AI, not a demo.
See the platform →
Strategy & Visualization
Thesis + Visual Explainers
Thesis: A multi-agent strategy platform where 21 specialized agents don't just answer — they discuss. Autonomous meeting rooms, Neo4j stakeholder graphs, Mem0 persistent memory. Visual explainers: 132 self-contained HTML pages — architecture diagrams, maturity models, process maps, audit reports — spanning every project. No frameworks, no build step, instant render.
Browse 132 explainers →
How I Work
Governance-first. Deterministic over LLM.
These aren't buzzwords. They're architectural commitments that appear in every platform.
Independent evaluation
- Agents evaluated by other agents, not themselves
- Programmatic scoring — no LLM self-assessment ever
- Weighted tiers + hard veto + certification levels
Audit everything
- Append-only logs — every decision traceable
- Row-level security at the database layer
- Human-in-the-loop gates on high-stakes actions
No vendor lock-in
- Provider abstraction — swap models without rewriting
- Automatic fallback on auth or rate-limit failures
- Anthropic for quality, DeepSeek for cost (~10× cheaper)
Ship with evidence
- 1,050+ automated tests across all platforms
- Structured JSONL logging on every backend
- Every platform has a visual explainer explaining it