Built with OpenAI

Custom OpenAI development, built for production.

KUMO designs and ships OpenAI agents, document workflows, and product features for revenue stage teams. We build with the Responses API, evals, model routing, cost controls, and full handover.

11Countries served
MetaTech Provider
AWSPartner

What KUMO builds

What a KUMO OpenAI build actually delivers

Agents

Agents and tools

Responses API workflows with tools, function calling, file search, web search, and stateful patterns. Function calls into your CRM and product APIs. Task specific validation and routing to human review. Model selection across the current OpenAI family per step.

Document AI

Document AI

Vision extraction, structured output with schema validation, exception queues, and human review for disagreements. Applicable to invoices, contracts, medical records, financial filings, and multi document research.

Product AI

Product AI

Search, classification, recommendations, support workflows, and embedded AI product features. Model routing, prompt engineering, and evaluation harness built into the product lifecycle rather than added on later.

Reliability & cost

Reliability and cost engineering

Evaluation harness, task based model routing, prompt caching, Batch API where latency permits (OpenAI documents 50 percent lower cost than synchronous APIs for eligible workloads), observability, and rollback. Cost attribution per user, per feature, per model.

When custom OpenAI is appropriate

When a custom OpenAI build makes sense

Four situations where a custom OpenAI build is worth the investment over an off the shelf product.

Document workflows at production scale

When your workflow ingests high volumes of unstructured or semi structured documents (invoices, contracts, medical records, financial filings, manifests) and needs structured output with validation, exception queues, and audit trails.

Agent workflows with real tool use

When your workflow needs multi step orchestration across your CRM, product APIs, internal systems, and third party services, with stateful sessions and traceable decisions. The current Responses API and tool interfaces support this pattern.

Product AI embedded in your software

When you are shipping AI features inside your own product (search, classification, recommendations, support workflows, in-product assistants) and need production controls, cost governance, and IP ownership rather than a black box vendor.

Workflows with material downside risk

When outputs carry material financial, legal, clinical, or customer risk, and you need evaluation gates, structured validation, and human review scoped to the risk tier. Not every AI decision needs the same control envelope.

Deployment decision

Direct OpenAI, Amazon Bedrock, or Azure OpenAI in Microsoft Foundry?

These are the routes most teams ask about. The right one depends on your procurement, region, identity, networking, retention, model availability, and cost profile. If your setup calls for a route that is not below, including running on your own infrastructure, we build for that too. The surrounding application runs in your cloud of choice.

Direct OpenAI

Direct OpenAI API

Best fit: fastest access to the latest models and features.

Consider when: your procurement supports direct commercial API, existing cloud is not a hard constraint, latency to OpenAI regions is acceptable.

AWS Bedrock

OpenAI models on Amazon Bedrock

Best fit: teams already on AWS, wanting first party OpenAI models under existing AWS commercial and data agreements.

Consider when: your compute, data, procurement, and identity are AWS aligned; region availability matches your needs.

Azure Foundry

Azure OpenAI in Microsoft Foundry

Best fit: teams already on Microsoft Azure, wanting Microsoft's data processing agreements and regional deployment.

Consider when: existing Microsoft cloud commitments, procurement requirements, or specific data residency requirements exist.

Whatever deployment your workflow needs, we build it. We compare identity, networking, procurement, residency, retention, latency, capability, and cost before selecting the route in the architecture review. Model and regional availability keeps moving, so we confirm the current catalog for your region rather than working from a fixed list.

Production operating model

How KUMO controls quality, cost, and risk

Every KUMO OpenAI engagement starts with a common control baseline. The specifics scale with your workflow's risk tier and error cost.

Evaluation gates

Versioned test sets, quality thresholds, regression checks, go/no-go criteria. No prompt or model change ships without an eval delta review.

Structured validation

Schemas, business rules, grounded citations where appropriate, and fallback paths for disagreements. Structured output enforced at the API layer.

Human control

Approval points for high impact outputs, clear escalation rules, and reviewer training. Risk based, not applied uniformly.

Cost and latency controls

Task based routing across the OpenAI family, prompt caching, Batch API for asynchronous jobs, budget alerts, and usage attribution per user and per feature.

Data and access controls

Data minimization, PII handling, retention policy, identity, secrets, network path, audit logs. Deployment route selected against these requirements.

Operations and handover

Dashboards, runbooks, rollback paths, incident ownership, source code, documentation. Handover designed so your team can operate the system without a mandatory retainer.

Not sure whether OpenAI is the right stack for your workflow?

Bring one workflow. Its inputs, its accuracy requirement, its cost ceiling, its latency budget, and its risk tier. Get a 30 minute fit call. KUMO tells you which OpenAI models fit, which deployment route makes sense, whether alternatives fit better, and what the production build looks like.

Book a consulting call

Delivery approach

How KUMO ships OpenAI builds

011 to 2 weeks · paid

Architecture review

Map workflow, inputs, accuracy requirement, cost ceiling, latency budget, and risk tier. Model selection against workflow specific evals. Deployment route decision. Application layer runs on your cloud of choice.

021 to 3 weeks

Data audit and eval design

Data quality review. Redaction and PII plan. Eval dataset built from real inputs. Success metrics defined per workflow. Human review checkpoint design.

032 to 4 weeks

Prototype and evals

Working prototype against real data. Eval harness live. Accuracy, safety behavior, and cost per query measured. Formal go or no go review against agreed success criteria.

043 to 8 weeks

Production build

Full production integration. Model routing across the OpenAI family. Prompt caching and Batch API where appropriate. Validation and fallback wired. Human review checkpoints deployed. Observability dashboards live.

051 to 2 weeks

Cost and observability tuning

Real cost per user and per feature measured. Caching and routing effectiveness verified. Model routing thresholds tuned against evals. Budget alerts wired to your on call rota.

061 to 2 weeks + optional

Handover and support

Runbooks handed over. Your team trained. Optional retainer for prompt tuning, model refresh when OpenAI ships new versions, and new workflow additions.

Typical engagement

What an OpenAI implementation costs

Three shapes of engagement based on scope and stage. OpenAI model usage is billed by the selected provider (direct OpenAI, Amazon Bedrock, or Azure OpenAI) under its current pricing. Final quote after the architecture review.

Project · Larger scope

Grow Build

Typical investment

$25K to $50K

Timeline: 10 to 20 weeks

Multi agent or multi workflow OpenAI system with model routing across tiers, deep tool integration via the Responses API, cross workflow observability, migration from a hosted chatbot platform, or compliance sensitive setup with Bedrock or Azure OpenAI.

Annual · For high volume business

Yearly Engagement

Typical investment

$50K to $100K / year

Timeline: Annual, ongoing partnership

Ongoing multi project engagement for teams running OpenAI at high volume. Continuous prompt tuning, model refresh when OpenAI ships new versions, new workflow additions, incident response, cost engineering, cross region expansion. Renews annually.

What changes the range: model mix across the OpenAI family, workflow complexity, integration depth, data volume, compliance and residency requirements, and how much of the eval harness is greenfield. OpenAI model usage is billed by the selected provider under its current pricing. Final quote after the architecture review.

Why KUMO fits

Why teams choose KUMO for production OpenAI work

Senior team, direct delivery

A senior product engineering team for 4 to 20 week production builds. The people who scope the workflow are the people who ship and operate it. Best for teams that need production engineering without a multi quarter transformation program.

Cost control by design

Task based model routing, prompt caching, and Batch API where latency permits are designed into the system. Before launch, we define cost per successful task, budget alerts, and the reporting your team needs.

Risk based evaluation and validation

Evaluation harness, structured validation, and human review scoped to each workflow's risk tier and error cost. Not every AI decision needs the same control envelope, and we design that envelope with you.

Full handover, source and runbooks yours

Full IP transfer at handover for the custom code, prompts, eval datasets, runbooks, dashboards, and rollback plans KUMO builds. Your team can operate the system without a mandatory retainer.

Related KUMO work

Where to go next

OpenAI and related resources

  • Case Studies KUMO's production AI and software work across fintech, SaaS, real estate, and more.
  • Built with Anthropic Claude Sibling page. Same delivery discipline, different model family.
  • Built with AWS Where most KUMO OpenAI builds run in production, including OpenAI models on Amazon Bedrock.

FAQ

OpenAI development FAQ

What is KUMO's edge on OpenAI builds?

Three things. A senior product engineering team that ships the system and stays through handover. Evaluation, validation, and cost controls designed in from day one, not added later. And KUMO operates its own production AI in CampaignHQ, so the patterns we ship for clients are the patterns we run for ourselves.

What replaces the OpenAI Assistants API?

OpenAI has scheduled the Assistants API for shutdown on 26 August 2026. Current builds use the Responses API with tools, function calling, file search, web search, and stateful workflows. If you have an existing Assistants API build, KUMO can migrate it to the Responses API before shutdown. See the next question.

Can KUMO migrate an existing Assistants API build to the Responses API before shutdown?

Yes. A migration typically includes: mapping current threads, runs, and tool calls to Responses primitives; rewriting orchestration and state handling; regression testing on your existing evaluation dataset; and cutover with rollback path. Scope depends on your current build size and integration depth. Architecture review confirms the scope and timeline for your specific migration.

Which OpenAI models does KUMO build with?

We build across the whole OpenAI family rather than committing to one model. OpenAI ships new versions often, so naming a favourite would date badly and serve you worse. In practice we route by task: a reasoning tier for the hardest steps, a balanced tier for most production work, and a cost tier for high volume paths. Which model sits in each slot is a decision we make against your evals during the architecture review, and we re-test on your eval harness whenever OpenAI ships a new version, so upgrades stay measurable and reversible. Existing deployments are maintained or migrated on measured quality, latency, support, and cost rather than on release notes.

How do OpenAI, Claude, and Gemini compare for our workflow?

Depends on the workflow. We evaluate candidates against your actual evaluation dataset in the architecture review rather than relying on marketing benchmarks. If another provider is a better fit for your accuracy target, latency budget, ecosystem, or procurement requirements, we say so. Model selection is a project decision, not a religion.

What does an OpenAI API bill look like at production scale?

Depends on model mix across the OpenAI family, volume, prompt caching effectiveness, and Batch API eligibility. Prompt caching can reduce cost meaningfully on cacheable workloads. OpenAI documents Batch API at 50 percent lower cost than synchronous APIs for eligible asynchronous workloads. We model your specific workload during architecture review with cache hit assumptions, model mix, and volume, and ship cost dashboards so you always know per user and per feature spend.

What about data privacy, retention, and Zero Data Retention?

OpenAI states that data sent to its commercial API is not used to train or improve models unless the customer opts in. That does not mean every endpoint has zero retention: abuse monitoring logs may be retained for up to 30 days, and the Responses API has 30 day application state retention by default. Eligible customers may apply for Modified Abuse Monitoring or Zero Data Retention. On Amazon Bedrock or Azure OpenAI, additional cloud provider agreements apply. KUMO designs data minimization, access, logging, retention, and human review controls for the selected deployment route.

How do you test quality before and after launch?

Every prompt or model change goes through a versioned evaluation dataset with quality thresholds and regression checks. Success metrics are defined per workflow during architecture review and revisited in the go/no-go review. Post launch, evals continue as regression tests when OpenAI ships new model versions or KUMO ships prompt or tool changes.

What about Codex, voice, and other OpenAI specialist products?

KUMO can evaluate current OpenAI speech, realtime, and Codex products for buyers whose workflows genuinely need them. Codex is currently positioned by OpenAI as a coding agent and software engineering product. The speech stack has moved on from the original transcription model to newer transcription and realtime options, and we pick against the current line up rather than an older default. Voice or Codex scope is a project decision made in architecture review, not a default part of every build.

What about IP ownership and handover?

Full IP transfer at handover for the custom code, prompts, eval datasets, runbooks, dashboards, and rollback plans KUMO builds. OpenAI's models remain OpenAI's. The system, your data, and the operating documentation are yours. Your team can operate without a mandatory KUMO retainer.

KUMO has applied to the OpenAI Partner Network. Application pending. KUMO does not claim a partner tier or specialization.

Ready to put an OpenAI workflow into production?

30 minutes, no pitch deck. You describe one workflow. Its inputs, its accuracy requirement, its cost ceiling, its latency budget, its risk tier. KUMO tells you which OpenAI models fit, which deployment route makes sense, whether alternatives fit better, and what the production build looks like.