Whitepaper/Supporting research

Scope and evidence

Client deployment, capacity and economics

Planning edition · 4 October 2026. Nike is a hypothetical client. These are engineering and commercial assumptions, not capacity already achieved, a vendor quotation or an agreed contract.

Who hosts and who pays

ArrangementInfrastructure and dataProvider billAgency role
Client cloud from outset: recommended enterprise pathNike-owned accounts, region and recordsNike's provider account pays directlyDesign/build, deployment, licensed software and agreed management; persona CPC where measurable
Agency-managed pilotIsolated client environment operated by usDirect client account preferred; otherwise funded/capped pass-throughSame services, with explicit usage funding and no unlimited allowance
Later handoverTransfer approved data, configuration and deployment to client accountsReplace credentials and billing with client-owned accountsDocumented transition; maintenance/licence rights agreed separately

Hosting the application on Nike servers does not move an OpenAI hosted model onto Nike GPUs. Approved content sent for hosted inference leaves Nike infrastructure; provider processing/retention must be reviewed separately. Entirely self-hosted inference is a different architecture requiring suitable model rights, hardware and voice/quality benchmarks. No such deployment has been built.

Plain-English request flow

The Nike app authenticates a member and asks a Nike-hosted session service for access. That service checks eligibility and budget, selects the approved Pre revision and authorises a limited session. Audio can travel directly between device and provider rather than through our application servers. Controlled tools retrieve permitted Nike knowledge; minimal usage events return to Nike's dashboard. Personal context is included only where needed and permitted.

Our local demo already uses direct WebRTC audio to OpenAI, with a server-side session handshake. It remains loopback-only and lacks production authentication, verified server metering, retrieval, logging and cost enforcement. Its five-minute browser timer is not a dependable financial cap.

Why memory is manageable

Share immutable persona versions and knowledge indexes; do not create a complete model or corpus copy per person. Keep small session records with an expiry and retrieve only relevant evidence. Durable relationship memory is a separate opt-in feature, not an unlimited transcript history.

An assumed 64 KB application record per active session means roughly 64 MB for 1,000 sessions or 640 MB for 10,000. These decimal estimates exclude application runtime, connections, replicated stores, caches and provider-held model context; they are not a measured memory profile. Provider context can still be expensive even when our own RAM is small.

At 100 average concurrent users and three minutes per conversation, a 30-day month produces 1.44 million sessions. An assumed 8 KB retained text record per session is about 11.52 GB/month before indexes, replicas and backups. Recording uncompressed mono audio at 16 kHz/16 bit would instead take 1.92 MB/minute; 4.32 million minutes would be about 8.29 TB. This audio format is an arithmetic illustration, not the WebRTC codec. Prefer temporary text and diagnostics; do not record audio by default.

Voice processing: a sourced estimate, not a fixed minute tariff

The present model is gpt-realtime-2.1. Published prices checked 4 October: audio input $32/million tokens; audio output $64/million; cached audio input $0.40/million. Text input/output are $4/$24 per million, cached text input $0.40. Exact model rate card.

For new speech, Realtime documents approximately 600 user audio tokens and 1,200 assistant audio tokens per spoken minute. Each reply also processes conversation context, with caching affecting cost; input transcription has a separate charge. Voice cost mechanics.

Assume each connected minute contains 30% user speech, 30% assistant speech and 40% no speech, with VAD filtering empty input. New speech alone costs (180 × $32 + 360 × $64) / 1,000,000 = $0.0288 per connected minute. This is not an all-in cost: repeated history, instructions, text output, transcription and tools are additional. gpt-4o-mini-transcribe's published estimated transcription rate is $0.003/audio minute. Pricing.

Use $0.05, $0.10 and $0.20 per connected minute as low/base/stress planning allowances for provider usage until measured. These are assumptions, not bounds: long, uncached or talk-heavy sessions can exceed them. Capture provider usage per reply and reconcile it against invoices; character count is not token count. Do not promise a saving merely by reducing stored source count.

Concurrency is not monthly volume

Monthly connected minutes = average concurrent sessions × 60 × 24 × 30. Peak sessions determine capacity and quotas; the average determines volume. These scenarios exclude cloud infrastructure, media, agency fees, taxes, data licences and evaluation calls.

ScenarioPeak / average concurrentMonthly minutesProvider allowance at $0.05 / $0.10 / $0.20
Limited pilot100 / 10432,000$21,600 / $43,200 / $86,400
Enterprise launch1,000 / 1004,320,000$216,000 / $432,000 / $864,000
Large rollout10,000 / 1,00043,200,000$2.16m / $4.32m / $8.64m
1,000 occupied continuously1,000 / 1,00043,200,000$2.16m / $4.32m / $8.64m

One million three-minute sessions is three million minutes: $150,000 / $300,000 / $600,000 under those allowances. The earlier $300,000 example was an assumption, not a measured provider charge or evidence that we outperform Meta/YouTube.

Cloud, operation and commercial costs

As a small compute illustration, four 1-vCPU/2-GB Linux x86 Fargate tasks running 720 hours cost approximately $142.19/month using AWS's US East example rates ($0.000011244/vCPU-second and $0.000001235/GB-second). This is CPU/RAM only, not a claim those tasks can support 1,000 conversations. Database, retrieval, network, load balancing, security, monitoring, regional redundancy and staff are extra. AWS pricing.

An initial $1,000–$10,000/month non-inference cloud allowance is a provisional budgeting envelope for scoping, not a quote for either load tier. Replace it with a region-specific architecture estimate after load tests. Enterprise support, data licences, rights clearances and evaluation need separate budgets.

Agency economics: design/build fee + deployment fee + agreed management/software fee + verified outward handoffs × persona CPC. Client economics include these plus platform entry-media spend, inference/hosting where applicable and client integration work. Performance billing never substitutes for paying the underlying delivery costs. No commission or revenue share from OpenAI is assumed.

Compare paid acquisition against incremental customer acquisition cost and contribution after returns, rather than dividing voice minutes by impressions. Compare member service against retention, repeat purchases and actual avoided service costs. Use controlled tests and avoid double-counting attribution. “The same Nike spirit. A relationship that's yours.” is the member proposition, not a guaranteed return.

Capacity and spending controls before launch

  1. Measure short, long, quiet, interrupted, failed and tool-heavy sessions. Establish p50/p95 latency, successful session rate, cost/minute and cost/useful outcome.
  2. Agree launch peaks, regions, sustained throughput and burst starts; secure provider quotas independently of cloud autoscaling. Validate token/request limits and simultaneous-session eligibility with the provider.
  3. Load-test application services with mock provider traffic, then perform a bounded authorised real end-to-end test. Scale 10 → 100 → 1,000 before a 10,000 target. Measure retrieval and usage-event load too.
  4. Authenticate sessions, rate-limit abuse, reserve a budget before accepting work, and enforce time/token/tool ceilings server-side. Reconcile reservations with actual provider charges and allow for delayed accounting and in-flight usage. A dashboard alert alone is not a hard stop.
  5. Define a capped fallback or graceful refusal at capacity; no retry storm. Keep operational events separate from sensitive transcripts and maintain incident/rollback procedures.

Build in the client cloud or hand over

The delivery package should include licensed source/artifacts as agreed, dependency inventory, deployment templates, secrets inventory without secret values, data schemas/export, approved persona/knowledge versions, Gauntlet cases and results, monitoring, backup/restore, incident and rollback runbooks, and operator training.

Acceptance includes deployment by Nike staff, credential rotation, access revocation, data transfer/deletion verification, a restore exercise and an agreed support window. Source transfer, platform licensing and ownership are distinct negotiated items; we cannot assume either Nike ownership of our whole platform or perpetual dependency on us.