Agents that finish the job
Software that plans a task, uses tools to carry it out and stops when it should — with step limits, spend caps, approval gates and a full trace of what it did.
We build AI features and the products around them — assistants that answer from your own documents, agents that complete real work, and the web applications they live in. Remote, worldwide, and delivered as working software rather than slide decks.
Most AI projects fail on architecture, not on the model. On whether a workflow or an agent fits the problem, where state lives, what happens when a tool errors, and who approves an action that cannot be undone.
Drawn in code, not photographed: every line here is SVG generated by the same build that produced this page. The screen types, the sketch draws itself, the city flickers.
Most teams arrive with a problem and a budget, not a spec. The first week is spent turning one into the other, and occasionally concluding you do not need a custom build at all.
Software that plans a task, uses tools to carry it out and stops when it should — with step limits, spend caps, approval gates and a full trace of what it did.
Workflow or agent, where retrieval sits, which model handles which path, what happens when a tool fails, and what it costs at ten times the volume. Decided first, written down, and testable.
The web application the AI lives inside: authentication, billing, admin tooling and an interface your team can operate without a manual.
A chat surface, an agent trace and an evaluation table, each assembled line by line in the order we would actually write them.
Judged on how many tickets resolve without a handoff, measured against your current rate.
Judged on handling time and how often a person has to correct the output.
Usually retrieval or chunking rather than the model. Judged on a before-and-after suite.
Permissions respected, every answer showing the source it came from.
Boring choices where boring is right, and newer tools only where they earn their place. Everything runs in accounts you own.
Nothing here is chosen because it is new. Each has to survive the same three questions: can you hire for it, will it still be maintained in three years, and how expensive is it to leave. Where your team already runs something reasonable, we use that instead.
Answer ten and you get a written brief you can copy into an email. It costs you two minutes and saves a week of back-and-forth. Nothing is sent anywhere until you send it.
These are the ten things we need to know before quoting. Answer them in an email if you prefer.
A project that should not exist is cheapest to stop in week one. If an off-the-shelf tool solves it, we will say so.
Six to ten weeks, start to handover. No discovery phase that ends in a slide deck.
Before any planning, we build the hardest screen with your real model behind it. It's faster to argue with something clickable.
Slow responses, empty results, a model that gets it wrong. These are the states users actually meet, so they get designed first.
Typed components, tested, in your repo and your design system. No handoff document — we write the code your team keeps.
A walkthrough with the engineers who inherit it, and two weeks on call after launch.
The five that come up before almost every engagement. The full list, including contracts and confidentiality, is on the FAQ page.
A scoped feature runs $8,000 to $35,000; a complete product $18,000 to $60,000; a small well-defined build starts at $2,500. The figure is fixed in writing before anything begins, and third-party service costs are billed at cost rather than marked up.
No. We work against whatever exists — a rough endpoint, a notebook, sometimes a person standing in for the model. Waiting for a stable API before designing the interface is how teams discover in month four that the interaction was wrong in month one.
Because the measure is agreed before the build: the metric, the baseline it is measured against, the method, and the bar the system has to clear. If those four cannot be settled, that is a signal the project is not ready to start.
Access to a working endpoint, one engineer we can ask questions of, and one person who can decide without a committee. That is genuinely it — research, states and edge cases come from our side.
You do, outright, on full payment — code, prompts, evaluation suites and documentation. Reusable tooling underneath it comes with a perpetual licence, and any open-source components are listed at handover so your team knows the obligations that came with them.
Two connected things: AI features (assistants, agents, retrieval systems, automations, document processing) and the software they need around them (web apps, APIs, dashboards, integrations). Most projects are a mix of both, because an AI feature with no product around it rarely reaches users.
Yes. Work is fully remote and clients are spread across North America, Europe, the Middle East, Australia and Asia. We keep a few hours of overlap with your working day, communicate in writing by default so nothing depends on a meeting, and deliver on your repository and your cloud account.
A working prototype against your real data usually takes one to two weeks. A production feature with authentication, monitoring, error handling and a proper hand-over is typically four to eight weeks, depending on how many systems it has to touch.
Small scoped builds start around $2,500. Most production projects land between $8,000 and $35,000. Monthly retainers for ongoing work start at $3,000. Every project gets a fixed number in writing before any work begins — see the pricing page for the detail.
Send a short description of the problem and we will reply within one business day with an honest view of scope, cost and whether we are the right person for it.
Or email directly: contact@hire-ai-dev.com