What is KUMO's edge on OpenAI builds? +
Three things. A senior product engineering team that ships the system and stays through handover. Evaluation, validation, and cost controls designed in from day one, not added later. And KUMO operates its own production AI in CampaignHQ, so the patterns we ship for clients are the patterns we run for ourselves.
What replaces the OpenAI Assistants API? +
OpenAI has scheduled the Assistants API for shutdown on 26 August 2026. Current builds use the Responses API with tools, function calling, file search, web search, and stateful workflows. If you have an existing Assistants API build, KUMO can migrate it to the Responses API before shutdown. See the next question.
Can KUMO migrate an existing Assistants API build to the Responses API before shutdown? +
Yes. A migration typically includes: mapping current threads, runs, and tool calls to Responses primitives; rewriting orchestration and state handling; regression testing on your existing evaluation dataset; and cutover with rollback path. Scope depends on your current build size and integration depth. Architecture review confirms the scope and timeline for your specific migration.
Which OpenAI models does KUMO build with? +
We build across the whole OpenAI family rather than committing to one model. OpenAI ships new versions often, so naming a favourite would date badly and serve you worse. In practice we route by task: a reasoning tier for the hardest steps, a balanced tier for most production work, and a cost tier for high volume paths. Which model sits in each slot is a decision we make against your evals during the architecture review, and we re-test on your eval harness whenever OpenAI ships a new version, so upgrades stay measurable and reversible. Existing deployments are maintained or migrated on measured quality, latency, support, and cost rather than on release notes.
How do OpenAI, Claude, and Gemini compare for our workflow? +
Depends on the workflow. We evaluate candidates against your actual evaluation dataset in the architecture review rather than relying on marketing benchmarks. If another provider is a better fit for your accuracy target, latency budget, ecosystem, or procurement requirements, we say so. Model selection is a project decision, not a religion.
What does an OpenAI API bill look like at production scale? +
Depends on model mix across the OpenAI family, volume, prompt caching effectiveness, and Batch API eligibility. Prompt caching can reduce cost meaningfully on cacheable workloads. OpenAI documents Batch API at 50 percent lower cost than synchronous APIs for eligible asynchronous workloads. We model your specific workload during architecture review with cache hit assumptions, model mix, and volume, and ship cost dashboards so you always know per user and per feature spend.
What about data privacy, retention, and Zero Data Retention? +
OpenAI states that data sent to its commercial API is not used to train or improve models unless the customer opts in. That does not mean every endpoint has zero retention: abuse monitoring logs may be retained for up to 30 days, and the Responses API has 30 day application state retention by default. Eligible customers may apply for Modified Abuse Monitoring or Zero Data Retention. On Amazon Bedrock or Azure OpenAI, additional cloud provider agreements apply. KUMO designs data minimization, access, logging, retention, and human review controls for the selected deployment route.
How do you test quality before and after launch? +
Every prompt or model change goes through a versioned evaluation dataset with quality thresholds and regression checks. Success metrics are defined per workflow during architecture review and revisited in the go/no-go review. Post launch, evals continue as regression tests when OpenAI ships new model versions or KUMO ships prompt or tool changes.
What about Codex, voice, and other OpenAI specialist products? +
KUMO can evaluate current OpenAI speech, realtime, and Codex products for buyers whose workflows genuinely need them. Codex is currently positioned by OpenAI as a coding agent and software engineering product. The speech stack has moved on from the original transcription model to newer transcription and realtime options, and we pick against the current line up rather than an older default. Voice or Codex scope is a project decision made in architecture review, not a default part of every build.
What about IP ownership and handover? +
Full IP transfer at handover for the custom code, prompts, eval datasets, runbooks, dashboards, and rollback plans KUMO builds. OpenAI's models remain OpenAI's. The system, your data, and the operating documentation are yours. Your team can operate without a mandatory KUMO retainer.
KUMO has applied to the OpenAI Partner Network. Application pending. KUMO does not claim a partner tier or specialization.