Total provider spend
–
AWS + OpenAI API · Same reporting period
The cost of running TollChat
What it costs to build, test, and run this reference implementation. AWS infrastructure and OpenAI API usage, together in one place.
–
AWS + OpenAI API · Same reporting period
–
Unblended cost
–
Organization-wide · Provider-reported
Last 30 completed days
Loading billing snapshot…
Daily costs in USD. Credits appear below zero. Billing may revise earlier days.
| Date | AWS | OpenAI | Total |
|---|
Month to date · Available AWS sources
EC2 – Other is AWS’s billing category for costs such as storage and data transfer. OpenAI API spend is shown separately above.
Month to date · Available AWS sources
Shared infrastructure supports the project across environments. Unallocated charges are included in the total; they have no recognized environment billing tag.
A different question
Inference is already part of provider spend. Per-answer attribution is not yet available.
Completed turns, model calls, and failed or retried calls need to be measured together. Missing usage is not free usage.
Infrastructure keeps the system available between chats. Dividing the entire cloud bill by a few conversations would not tell you the cost of an additional answer.
AWS uses Cost Explorer unblended cost for this deployment’s account. The production dashboard also reads development’s published AWS aggregate. OpenAI uses its organization Costs API, without project filtering. A combined total requires every source to cover matching UTC dates in USD. Development combines its AWS account with organization-wide OpenAI costs, which include all environments. Production counts OpenAI once, directly from the provider. AWS-hosted model charges remain within AWS. Token estimates are never added again.
All spend in these accounts and this OpenAI organization is treated as TollChat spend, including development and evaluations. No returned charges are excluded. Credits and adjustments retain their sign. AWS uses the lowercase environment tag: production, development, or shared; other charges are unallocated. OpenAI workload allocation is not established.
Queries cover the union of month to date and the last 30 completed days, ending at midnight UTC today. On the first of a month, month to date has no completed days. Requested dates are not a guarantee that providers have finished reporting them. Neither source confirms a finalized-through date here. Development refreshes at 08:00 UTC and production at 09:00 UTC daily. Failed refreshes retain the last valid snapshot; publication older than 48 hours is flagged.
For each controlled scenario: attributed inference cost, including measured failed calls and retries, divided by completed assistant turns. A turn runs from one submitted user message to its terminal answer and may use multiple model calls. Usage, verified rate sources, effective dates, and non-overlapping token categories are required before any estimate is shown.