01
Production-ready LLM / RAG / agent application
The actual product, not a notebook. Typed API surface, background jobs where they belong, and failure modes handled instead of ignored.
PRODUCT 01
3–5 weeks
Production agents, RAG pipelines and eval harnesses that your customers actually use — built by a studio that builds with agents.
WHO THIS IS FOR
Founders with an AI idea or working prototype who need it live and reliable.
Teams that have a demo but need production-grade reliability, evaluation, and observability.
Companies that want a real AI feature — not a chatbot toy — shipped in weeks.
WHAT YOU GET
01
The actual product, not a notebook. Typed API surface, background jobs where they belong, and failure modes handled instead of ignored.
02
Chunking, embedding and indexing on pgvector or an equivalent store, with hybrid keyword + vector search and reranking where it measurably helps.
03
A golden dataset and repeatable eval runs so model quality is a number you can watch, not a vibe. Regressions get caught before your users find them.
04
Sign-in, sessions, roles, and per-tenant data isolation if the product needs it.
05
Structured request and trace logging, prompt/response capture, latency and error monitoring, and per-model token and cost tracking.
06
For every task the system performs autonomously: what it may do, what it must escalate, what it may never touch. The interesting engineering in an agent product is the edges, and they belong in documentation rather than in someone's head.
07
Dockerised services, environment and secret handling, CI, and a deploy you can run yourself.
08
Architecture notes, runbook, and a live walkthrough with whoever will own the code next.
09
Something breaks in the first fortnight, we fix it.
WHAT IT LOOKS LIKE
| retrieval@5 | 0.94 | +0.02 |
| answer_grounded | 0.98 | +0.01 |
| refusal_correct | 0.91 | −0.03 |
| citation_resolves | 1.00 | 0.00 |
| p95_latency_ms | 1,840 | −260 |
refusal_correct regressed — build blocked, not shipped. Quality is a number we watch, not a claim we make.
ceiling €0.50
budget 2s
cap €80 → queues
tuned weekly
Budgets are code, not a monthly surprise. Crossing the per-run ceiling switches to a cheaper path; crossing the daily cap queues and alerts.
HOW IT WORKS
01
One call and one document. We agree exactly what ships, what the success criteria are, and what is explicitly out of scope.
02
Model and retrieval strategy, data flow, tenancy, and the eval set that defines 'good'. Decisions written down before code.
03
Weekly deployable increments against your real data. Evals run every iteration, so quality moves in one direction.
04
Rate limits, cost ceilings, monitoring, load sanity checks, production deploy, documentation and a training session.
TIMELINE & PRICING
DURATION
3–5 weeks
BUDGET RANGE
€6k – €30k+
Typically €10k – €16k
PAYMENT
40% to start · 40% at midpoint · 20% on delivery.
WHERE IN THE RANGE YOU LAND
BEYOND THE PRODUCT
Multi-team platform builds, long-running agent systems and anything with a compliance programme attached run €50k – €200k+ and are scoped as an engagement rather than a product. Same engineering, different shape: monthly phases, a named team, and a roadmap instead of a fixed scope.
STACK
NOT INCLUDED
FAQ
BUDGET · €6k – €30k+
Tell us what you're building and what 'working' looks like. We'll come back with scope, timeline and a fixed price.
Start this project →