Cortex
Internal tooling for my own workflow: a local-first multi-agent system that ingests from several sources, classifies and enriches with LLMs through a tiered routing policy, and stores validated structured output in a vector database, encrypted at rest.
Context
A personal system for turning a stream of incoming material (RSS and news, finance data, meeting and message sources) into classified records and structured insights I can query later.
Problem
The problems are the ones any pipeline has: how to spend inference budget sensibly, how to get output that downstream code can rely on, and how to store personal data safely. A system that calls the best available model for every task and parses prose out of the response is fine to demo and unreliable once it runs unattended on a schedule.
What I did
- Split the system across two hosts, a 24 GB VM serving models locally and a 4 GB VM running orchestration, connected over gRPC with streaming completions and an embedding endpoint.
- Built a tiered model-routing policy: tasks declare a tier by importance, and the registry resolves tier plus domain to a specific model, with fallbacks.
- Layered the pipeline as ingestion → classification → orchestration → LLM → storage, so that cheap deterministic work happens before anything reaches a model.
- Required validated structured output at every LLM boundary: responses are parsed into a strict schema, normalized, and rejected instead of stored when they don't conform.
- Encrypted payloads at rest with Fernet before they are written to the vector store, for both raw records and generated insights.
- Exposed the result through a GraphQL API over the stored records.
Two hosts
The model process needs 24 GB of memory and is busy for seconds at a time. The orchestration service needs predictable availability, does mostly IO, and gets restarted whenever I change something. Running both in one process means every model load competes with the part that is supposed to stay up.
So they are separate hosts: a 24 GB VM running the model server and registry, and a 4 GB VM running ingestion, classification, storage and the API. They talk over gRPC, with completions streamed token by token and a separate endpoint for embeddings.
Model routing
Sending everything to the strongest available model is expensive where it is not slow. Deciding which category a record belongs to is a smaller problem than writing a daily summary across everything that arrived.
So tasks declare a tier by how much quality matters: validation and classification at the low tier, insight generation in the middle, strategic and summary work at the top. Routing resolves tier *and* domain to a concrete model, because the model that handles finance text best is not the one that classifies best.
TIER_MODEL_MAP = {
"classification": {
"tier2": "llama3", # fast, high volume, low complexity
"tier3": "mistral", # better context handling
"tier4": "gpt-4", # only where the decision matters most
},
"finance": {"tier2": "mistral", "tier3": "mistral", "tier4": "mistral"},
"daily_note": {"tier2": "llama3", "tier3": "mistral", "tier4": "gpt-4"},
"default": {"tier2": "llama3", "tier3": "llama3", "tier4": "mistral"},
}The pipeline is tiered the same way: deterministic rules classify what they can, and the model is consulted only for what needs judgment. That also keeps the volume of model calls down on a host serving them locally.
The validation boundary
An LLM that returns prose is unusable as a component here. Once downstream code has to interpret free text, the parse can fail in ways nobody enumerated, and it fails quietly, storing a record that reads fine and carries the wrong values.
So every LLM boundary produces structured output that is parsed, validated against a schema, and normalized before anything is written. A response that does not conform is rejected instead of coerced into shape, because coercion produces records whose fields are populated and wrong.
{
"domain": "finance",
"priority": "high",
"sentiment": "negative",
"summary": "Concise generated statement.",
"source_ids": ["raw-record-uuid"],
"model": "mistral",
"tier": "tier3"
}Storage runs in the same order every time: payloads are encrypted with Fernet before they are written, for raw records and generated insights alike, and decrypted only when something reads them. This is personal data on machines I administer myself.
Result
- A working local-first pipeline: ingest, classify, enrich through a routed model call, validate, encrypt, and store as queryable vectors.
- Model choice is a routing-table entry, so swapping or adding a model does not touch call sites.
- Inference is isolated on its own host, so the memory-hungry half can be restarted or re-hosted without disturbing the half that holds the data.
- Structured, schema-validated output at every model boundary, which is what made the system dependable enough to run unattended.
- A solo side project on two VMs I administer myself. There is no production scale behind it, no users besides me, and no uptime claim.
Tech
- Python
- Qdrant
- Ollama
- llama.cpp
- gRPC
- LangGraph / CrewAI
- GraphQL
- Fernet