Engineering AI that Survives Production
Most AI features work in the demo and quietly fail in production. We build the parts that decide which one you get — the evals, the retrieval, the context — for US companies, inside US business hours.
- ~90%
- on a client's own benchmark
- 50×
- faster retrieval
- UTC−3
- US business hours
The Demo Gap
Why do AI features that demo perfectly die in production?
It is almost never the model. The gap between a demo that lands and a feature that holds is made of three things, and they are all boring.
Why it happens
A demo runs on inputs somebody chose
Clean, representative, hand-picked examples. Production sends everything else — the malformed, the adversarial, the merely unusual. The feature did not regress when it hit real traffic. It was never measured against real traffic in the first place.
So the failure feels sudden, when it was actually invisible from the start.
What has to change
You cannot improve what you never measured
Before any prompt change, model upgrade or retrieval tweak means anything, there has to be a measurement built from your actual failures. Without one, every change is a guess, and half of them make the system quietly worse.
With one, you find out a change hurt before your users do.
How we do it
Evals first, then context — not a bigger model
We build the eval from your real failures, then engineer what goes into the context window and why. Reaching for a larger model is the expensive way to buy an improvement you cannot verify and did not need.
One client's agents went from failing their own benchmark to roughly 90% on it, with no model change.
What we build
Four disciplines, each with a number attached.
AI & LLM systems
Agents that survive real inputs.
Agent architecture, retrieval, evals, MCP.
How Cosmos Reader's agents went from failing their own benchmark to ~90%
~90%Full-stack product
Systems that hold at load.
Python and TypeScript, front to back, plus the cloud bill.
How CoinPanel held 60,000 concurrent trades without degrading
60,000Shopify & Hydrogen
Storefronts that speak every market.
Headless commerce, internationalization, personalization.
How e.l.f. Beauty launched a storefront that speaks every market
Multi-marketAI digital marketing
Attention, then revenue — both measured.
Content pipelines wired to conversion machinery.
How Apex built one system for making attention and converting it
Both ends
Each links to how we work in that discipline
Tools
What we build with.
- Python
- TypeScript
- Node.js
- Kotlin
- Swift
- Linux
- FastAPI
- PostgreSQL
- Airflow
- Supabase
- Firebase
- Firestore
- Vercel
- AWS
- Google Cloud
- Kubernetes
- Docker
- Terraform
- React
- Next.js
- Remix
- React Native
- Expo
- Tailwind
- GSAP
- Shopify
- Hydrogen
- Builder.io
- Claude
- OpenAI
- Gemini
- Llama
- Qwen
- Cohere
- MCP
- Ollama
- Replicate
- FLUX
- ComfyUI
- SDXL
- Midjourney
- fal.ai
- Kling
- Runway
- Luma
- Seedance
- Wan 2
- MiniMax
- Veo 3
- Sora
- ElevenLabs
- Whisper
Receipts
Trusted by
Seventeen companies have shipped with us.
Toptal
Senior full-stack engineer · Top 3%
BCG
Forward deployed contractor
e.l.f. Beauty
Shopify storefront engineer
Apex
Founder & AI engineer
Fundamental Labs
Full-stack developer
Cosmos Reader
Senior full-stack engineer
The Form Factory
Commerce delivery
Tavern Research
AI back-end engineer
CoinPanel
Senior full-stack engineer
Cielo24
Senior back-end developer
Verusen
Senior DevOps & back-end
Retail Brand
Founder & full-stack engineer · NDA
ReliableSite
Back-end developer
Brick Abode
Senior software engineer
Covario
Back-end developer · via Brick Abode
BTC SKR
Full-stack developer
Trading For Liberty
Full-stack developer
How we work with you
In your hours, in your repo.
Florianópolis is UTC−3 — one to two hours ahead of New York, four to five ahead of San Francisco. A question at 10am is answered at 10am. Most of what goes wrong with distributed engineering is a scheduling problem in a technical costume: a wrong assumption survives twelve hours because nobody could challenge it in time.
- We join your standups, your repo, your board — no parallel process to reconcile later.
- Senior only. The person who hits the ambiguity is the person who can resolve it.
- Every engagement on this page has a number attached. That is the standard.
- San Francisco
- your morning
- New York
- your desk
- Florianópolis
- us, working
- 01
A call, not a questionnaire
Thirty minutes on what you are building and what is in the way. If we are not the right shop, we will say so on that call.
- 02
A scoped first slice
One real deliverable with a defined edge — an eval suite, a storefront section, a pipeline that was too slow. Small enough to judge, big enough to matter.
- 03
Embedded delivery
In your tools, your review process, your standups. Because we work your hours, review cycles close the same day.
- 04
Expand or stop, on evidence
The first slice tells you what we are worth. Scale up, or walk away having spent very little finding out.
Who you actually get
A small senior firm, not a bench.
Blueish Flower is led by Natan Votre — seven years across AI, backend and infrastructure, with 300+ engineers interviewed as a technical screener. Currently engaged with BCG, e.l.f. Beauty and Apex. You will meet the person who writes the code on the first call, and they will still be on the project at the end of it.
- 7+ years across AI, backend and infrastructure
- B.Eng. Electronics — Federal University of Santa Catarina
- Florianópolis, Brazil · UTC−3
Questions
The things buyers ask first.
Our AI feature works in demos but not in production. Can you fix that?
That is the most common engagement we take. The gap is almost always inputs: demos run on clean examples someone selected, and production does not. The work starts with an eval built from your real failures, because without one every change afterwards is a guess.
Are you tied to a particular model provider?
No, and we build so you aren't either. The routing and eval layers sit above the model so it can be swapped when a better or cheaper one ships — which, at the current pace, it will before your project is a year old.
How much of the workday do you overlap with US teams?
All of it. We work from Brazil at UTC−3, one to two hours ahead of US Eastern and four to five ahead of Pacific depending on daylight saving. A full US business day is a full business day for us — no night shift, no handoff document.
Is this offshore outsourcing?
It is nearshore, and the difference is the calendar. Offshore usually means a twelve-hour gap and asynchronous handoffs. We are in your meetings and your repo while you are at your desk.
What else do you build besides AI?
Full-stack product engineering, Shopify and Hydrogen commerce, and AI-assisted digital marketing. Four disciplines, deep rather than a menu of twenty. If a project sits outside them we will tell you instead of learning on your budget.
Who have you built for?
Seventeen companies, including BCG, e.l.f. Beauty, Toptal, CoinPanel, Cielo24, Verusen and Tavern Research — consultancies, agencies, startups and trading desks. One current retail client is under NDA and is listed unnamed.
How senior are the engineers who do the work?
Senior only, and we have interviewed over 300 engineers as technical screeners since 2022 — so the bar is not a claim, it is a job we have done from the other side of the table.
How fast can you start?
Days, not quarters. The first step is a thirty-minute call and a scoped first slice of real work, so there is no discovery phase to fund before anything ships.
What happens if the engagement is not working?
You stop. We start with one small, judgeable deliverable precisely so that decision costs you very little and arrives early, rather than surfacing three months into a fixed contract.
Do you sign NDAs?
Yes, routinely. One of our current engagements is under NDA and appears on this site without the client's name for that reason.


