Skip to content

Engineering AI that Survives Production

Most AI features work in the demo and quietly fail in production. We build the parts that decide which one you get — the evals, the retrieval, the context — for US companies, inside US business hours.

~90%
on a client's own benchmark
50×
faster retrieval
UTC−3
US business hours

The Demo Gap

It is almost never the model. The gap between a demo that lands and a feature that holds is made of three things, and they are all boring.

  1. Why it happens

    A demo runs on inputs somebody chose

    Clean, representative, hand-picked examples. Production sends everything else — the malformed, the adversarial, the merely unusual. The feature did not regress when it hit real traffic. It was never measured against real traffic in the first place.

    So the failure feels sudden, when it was actually invisible from the start.

  2. What has to change

    You cannot improve what you never measured

    Before any prompt change, model upgrade or retrieval tweak means anything, there has to be a measurement built from your actual failures. Without one, every change is a guess, and half of them make the system quietly worse.

    With one, you find out a change hurt before your users do.

  3. How we do it

    Evals first, then context — not a bigger model

    We build the eval from your real failures, then engineer what goes into the context window and why. Reaching for a larger model is the expensive way to buy an improvement you cannot verify and did not need.

    One client's agents went from failing their own benchmark to roughly 90% on it, with no model change.

Tools

What we build with.

  • Python
  • TypeScript
  • Node.js
  • Kotlin
  • Swift
  • Linux
  • FastAPI
  • PostgreSQL
  • Airflow
  • Supabase
  • Firebase
  • Firestore
  • Vercel
  • AWS
  • Google Cloud
  • Kubernetes
  • Docker
  • Terraform
  • React
  • Next.js
  • Remix
  • React Native
  • Expo
  • Tailwind
  • GSAP
  • Shopify
  • Hydrogen
  • Builder.io
  • Claude
  • OpenAI
  • Gemini
  • Llama
  • Qwen
  • Cohere
  • MCP
  • Ollama
  • Replicate
  • FLUX
  • ComfyUI
  • SDXL
  • Midjourney
  • fal.ai
  • Kling
  • Runway
  • Luma
  • Seedance
  • Wan 2
  • MiniMax
  • Veo 3
  • Sora
  • ElevenLabs
  • Whisper

Receipts

17
Companies shipped for
startups → enterprise
7+
Years in production
AI · backend · infra
60k
Concurrent trades held
trading platform
50×
Faster retrieval
parallelized RAG
$30k
Cloud spend removed
per month · Kubernetes rework
300+
Engineers vetted
our own hiring bar

Trusted by

How we work with you

Florianópolis is UTC−3 — one to two hours ahead of New York, four to five ahead of San Francisco. A question at 10am is answered at 10am. Most of what goes wrong with distributed engineering is a scheduling problem in a technical costume: a wrong assumption survives twelve hours because nobody could challenge it in time.

  • We join your standups, your repo, your board — no parallel process to reconcile later.
  • Senior only. The person who hits the ambiguity is the person who can resolve it.
  • Every engagement on this page has a number attached. That is the standard.
 
San Francisco
your morning
 
New York
your desk
 
Florianópolis
us, working
  1. 01

    A call, not a questionnaire

    Thirty minutes on what you are building and what is in the way. If we are not the right shop, we will say so on that call.

  2. 02

    A scoped first slice

    One real deliverable with a defined edge — an eval suite, a storefront section, a pipeline that was too slow. Small enough to judge, big enough to matter.

  3. 03

    Embedded delivery

    In your tools, your review process, your standups. Because we work your hours, review cycles close the same day.

  4. 04

    Expand or stop, on evidence

    The first slice tells you what we are worth. Scale up, or walk away having spent very little finding out.

Who you actually get

A small senior firm, not a bench.

Blueish Flower is led by Natan Votre — seven years across AI, backend and infrastructure, with 300+ engineers interviewed as a technical screener. Currently engaged with BCG, e.l.f. Beauty and Apex. You will meet the person who writes the code on the first call, and they will still be on the project at the end of it.

  • 7+ years across AI, backend and infrastructure
  • B.Eng. Electronics — Federal University of Santa Catarina
  • Florianópolis, Brazil · UTC−3

Questions

The things buyers ask first.

Our AI feature works in demos but not in production. Can you fix that?

That is the most common engagement we take. The gap is almost always inputs: demos run on clean examples someone selected, and production does not. The work starts with an eval built from your real failures, because without one every change afterwards is a guess.

Are you tied to a particular model provider?

No, and we build so you aren't either. The routing and eval layers sit above the model so it can be swapped when a better or cheaper one ships — which, at the current pace, it will before your project is a year old.

How much of the workday do you overlap with US teams?

All of it. We work from Brazil at UTC−3, one to two hours ahead of US Eastern and four to five ahead of Pacific depending on daylight saving. A full US business day is a full business day for us — no night shift, no handoff document.

Is this offshore outsourcing?

It is nearshore, and the difference is the calendar. Offshore usually means a twelve-hour gap and asynchronous handoffs. We are in your meetings and your repo while you are at your desk.

What else do you build besides AI?

Full-stack product engineering, Shopify and Hydrogen commerce, and AI-assisted digital marketing. Four disciplines, deep rather than a menu of twenty. If a project sits outside them we will tell you instead of learning on your budget.

Who have you built for?

Seventeen companies, including BCG, e.l.f. Beauty, Toptal, CoinPanel, Cielo24, Verusen and Tavern Research — consultancies, agencies, startups and trading desks. One current retail client is under NDA and is listed unnamed.

How senior are the engineers who do the work?

Senior only, and we have interviewed over 300 engineers as technical screeners since 2022 — so the bar is not a claim, it is a job we have done from the other side of the table.

How fast can you start?

Days, not quarters. The first step is a thirty-minute call and a scoped first slice of real work, so there is no discovery phase to fund before anything ships.

What happens if the engagement is not working?

You stop. We start with one small, judgeable deliverable precisely so that decision costs you very little and arrives early, rather than surfacing three months into a fixed contract.

Do you sign NDAs?

Yes, routinely. One of our current engagements is under NDA and appears on this site without the client's name for that reason.