EliconStart a project →
← ALL SERVICES

PRODUCT 01

3–5 weeks

AI Product Launch

Production agents, RAG pipelines and eval harnesses that your customers actually use — built by a studio that builds with agents.

Start this project →€6k – €30k+typically €10k – €16k
  • python
  • fastapi
  • openai
  • anthropic
  • langgraph
  • pgvector
  • postgres
  • redis
  • docker

WHO THIS IS FOR

Built for you if…

  • 01

    Founders with an AI idea or working prototype who need it live and reliable.

  • 02

    Teams that have a demo but need production-grade reliability, evaluation, and observability.

  • 03

    Companies that want a real AI feature — not a chatbot toy — shipped in weeks.

WHAT YOU GET

Everything included.

01

Production-ready LLM / RAG / agent application

The actual product, not a notebook. Typed API surface, background jobs where they belong, and failure modes handled instead of ignored.

02

Proper retrieval pipeline

Chunking, embedding and indexing on pgvector or an equivalent store, with hybrid keyword + vector search and reranking where it measurably helps.

03

Evaluation harness

A golden dataset and repeatable eval runs so model quality is a number you can watch, not a vibe. Regressions get caught before your users find them.

04

Authentication and multi-tenancy

Sign-in, sessions, roles, and per-tenant data isolation if the product needs it.

05

Observability

Structured request and trace logging, prompt/response capture, latency and error monitoring, and per-model token and cost tracking.

06

Agent boundaries written down

For every task the system performs autonomously: what it may do, what it must escalate, what it may never touch. The interesting engineering in an agent product is the edges, and they belong in documentation rather than in someone's head.

07

Deployment to a production environment

Dockerised services, environment and secret handling, CI, and a deploy you can run yourself.

08

Handover documentation + training session

Architecture notes, runbook, and a live walkthrough with whoever will own the code next.

09

2 weeks of post-launch critical bug support

Something breaks in the first fortnight, we fix it.

WHAT IT LOOKS LIKE

Some of what you get back.

EVAL RUN — 500 GOLDEN CASESrun 47 · vs run 46
retrieval@50.94+0.02
answer_grounded0.98+0.01
refusal_correct0.91−0.03
citation_resolves1.000.00
p95_latency_ms1,840−260

refusal_correct regressed — build blocked, not shipped. Quality is a number we watch, not a claim we make.

COST & LATENCY — ENFORCED IN CODElast 24h
  • cost / run€0.11

    ceiling €0.50

  • p50 latency740ms

    budget 2s

  • daily spend€38

    cap €80 → queues

  • escalation rate6.2%

    tuned weekly

Budgets are code, not a monthly surprise. Crossing the per-run ceiling switches to a cheaper path; crossing the daily cap queues and alerts.

HOW IT WORKS

The process.

  1. 01

    Discovery call + scope lock

    1–2 days

    One call and one document. We agree exactly what ships, what the success criteria are, and what is explicitly out of scope.

  2. 02

    Architecture & data/eval design

    Week 1

    Model and retrieval strategy, data flow, tenancy, and the eval set that defines 'good'. Decisions written down before code.

  3. 03

    Build + iterative testing with real data

    Weeks 1–4

    Weekly deployable increments against your real data. Evals run every iteration, so quality moves in one direction.

  4. 04

    Hardening, deployment, handover

    Final week

    Rate limits, cost ceilings, monitoring, load sanity checks, production deploy, documentation and a training session.

TIMELINE & PRICING

What it costs.

DURATION

3–5 weeks

BUDGET RANGE

€6k – €30k+

Typically €10k – €16k

PAYMENT

40% to start · 40% at midpoint · 20% on delivery.

WHERE IN THE RANGE YOU LAND

  • ·€6k–10k: one clearly defined feature on data you already have, single tenant, one integration.
  • ·€10k–16k: where most of these land — retrieval over your own corpus, evals, auth, observability, production deploy.
  • ·€16k–30k: multi-agent orchestration, several system integrations, or strict compliance and audit requirements.
  • ·€30k+: the ceiling is the scope, not the price list. Big systems are quoted after discovery.
  • ·Fixed price once scope is locked. We quote a number, not a rate card — you are buying a shipped product, not our hours.
  • ·Model, infrastructure and third-party API costs are billed to your accounts, not ours.

BEYOND THE PRODUCT

Multi-team platform builds, long-running agent systems and anything with a compliance programme attached run €50k – €200k+ and are scoped as an engagement rather than a product. Same engineering, different shape: monthly phases, a named team, and a roadmap instead of a fixed scope.

STACK

What we build it with.

  • Python
  • FastAPI
  • OpenAI
  • Anthropic
  • LangGraph
  • pgvector
  • Postgres
  • Redis
  • Docker

NOT INCLUDED

Where the line is.

  • ×Training or fine-tuning foundation models from scratch.
  • ×Model, cloud and third-party API spend.
  • ×Brand identity, marketing sites, or content production.
  • ×Ongoing 24/7 on-call beyond the support window (available separately as a retainer).
  • ×Formal compliance certification (SOC 2, HIPAA audits) — we build to sane practices, auditors are their own project.

FAQ

Questions we get.

We already have a prototype. Does that make it cheaper?
Usually faster rather than cheaper. A working prototype removes product ambiguity, which is the expensive part. Prototype code itself is often replaced — we keep what earns its place.
Which model should we use?
Whichever wins on your eval set at an acceptable cost and latency. We build the code so the model is a swappable choice, then measure instead of arguing.
How do you stop the thing from hallucinating?
Grounded retrieval, structured outputs with schema validation, refusal paths when confidence is low, and an eval harness that measures how often it happens. You get a number, not a promise.
Who owns the code?
You do, in full, from day one. Work happens in your repository where possible, and the handover includes everything needed to run it without us.
Can our own developers take it over afterwards?
That is the intended outcome. Conventional stack, documented architecture, and a training session with your team as part of delivery.
What does the agent actually own, and what stays human?
We answer that in writing during scoping, per task. The default is that agents own the repetitive, verifiable work and escalate anything with an irreversible consequence — and every autonomous action is logged well enough to audit after the fact.
You say you build with agents. Does that mean nobody reads the code?
The opposite. Agents write the first pass against tests; a senior engineer reads every line before it ships and is accountable for it at handover. The speed comes from the first pass, not from skipping review.
What if it turns out the AI part doesn't work?
We find that out in week one during eval design, not in week five. If the data cannot support the feature, you get a written explanation and a cheaper alternative rather than an expensive lesson.

BUDGET · €6k – €30k+

Have an AI product to ship?

Tell us what you're building and what 'working' looks like. We'll come back with scope, timeline and a fixed price.

Start this project →