There is no reliable universal price for “an AI agent.” A useful estimate names the workflow, degree of autonomy, integrations, quality target, risk controls, and operating period. Build estimates below are a worksheet structure, not market-rate guidance. Use your own compensation, vendor quotes, and project evidence.
Scope tiers
Prototype: one narrow task, test data, manual review, and no production write access. Estimate discovery, a thin integration, representative evaluations, and removal of test data.
Bounded production workflow: defined inputs and outputs, authenticated connections, explicit permission boundaries, monitoring, fallback behavior, and an accountable operator. Include security review and release work as named tasks.
Multi-workflow or higher-impact system: multiple integrations or user groups, durable operational support, formal change control, incident response, and ongoing evaluation. Break it into independently testable work packages; do not extrapolate a prototype estimate by multiplying by a guessed factor.
These are planning categories, not certified maturity levels or compliance classifications. Whether a system is suitable for deployment depends on its actual context and review.
Drivers
Estimate effort from requirements that can change the work: number and quality of data sources; permission and identity design; API stability; workflow exceptions; human review; evaluation data; latency and availability targets; auditability; deployment environment; and who owns incidents. A read-only draft assistant and a system that can modify customer records have different integration, testing, and approval work.
Price model usage from representative measured runs, including long-context cases, tool results, retries, and failed or abandoned workflows where billing rules apply. Check the relevant provider’s current pricing and billing documentation, such as OpenAI API pricing or Google Agent Platform pricing. Do not reuse a token rate across providers or infer an invoice from token charges alone.
Build vs buy
Compare options against the same workflow and period. For a product or service quote, record included usage, overages, implementation, support, data retention, security terms, portability, renewal, and exit. For an internal build, estimate discovery, integration, evaluation, release, and operational ownership. A vendor subscription can still require internal integration and oversight; a custom build can still depend on paid model and infrastructure services.
Use total cost and an outcome measure together. The FinOps Foundation’s Unit Economics guidance recommends connecting technology costs with a meaningful business unit and value measure; its framework is flexible rather than a prescribed vendor comparison. Define a shared denominator—such as completed, accepted tasks—and record quality and escalation alongside cost. Do not rank vendors from incomparable list prices.
Worked hypothetical comparison: an internal pilot estimate has 90 hours across discovery (20), integration (40), and evaluation/release review (30). At a fictional loaded rate of $80/hour, labor is 90 × $80 = $7,200. Add a fictional $600 in one-time external setup and $300/month for model, tools, and hosting over a six-month comparison period: $7,200 + $600 + (6 × $300) = $9,600. A hypothetical vendor quote of $1,200/month plus $1,500 setup would total $8,700 over the same six months before internal integration or oversight. Those sample amounts are invented arithmetic only; they do not describe market prices or an actual vendor offer. Add missing internal effort to both options before deciding.
Typical team
Assign responsibilities, even if one person covers several: workflow owner for acceptance criteria; engineering owner for integration and reliability; security/privacy reviewer for data flow and permissions; operations owner for monitoring and incidents; and finance/procurement for cost scope and commercial terms. Add domain experts and human reviewers where the workflow needs them. Estimate each role in hours rather than assuming a fixed team size.
For employee labor, use organization-specific loaded rates. The U.S. Bureau of Labor Statistics publishes occupation wage data and employer compensation data (see OEWS tables and Employer Costs for Employee Compensation), but national averages do not represent a particular team, location, seniority mix, contractor rate, or project quote. Use them only as broad context when local data is unavailable and state that limitation.
Ongoing run costs
Budget recurring model and tool consumption, hosting, logs and traces, evaluations, support, permission maintenance, vendor renewals, incident response, and periodic changes. Forecast low, expected, and high workload cases from observed or explicitly assumed task volumes and step counts. Track cost per accepted task alongside retries and human escalation; lower token cost can still be worse if quality problems create rework.
Separate one-time implementation from monthly operating cost, and record exclusions such as taxes, internal shared services, or a vendor’s minimum commitment. Revisit the estimate after a representative production billing period and whenever the model, workflow, contract, or volume changes. The AI agent cost calculator models user-entered run-cost assumptions; the LLM API cost calculator and token cost calculator help isolate token scenarios. They do not estimate engineering labor or retrieve live prices.