Capacity forecast
1–2 weeks
Translate product traffic and AI workflows into a defensible model-usage and cost range.
Estma service / Model capacity
Plan and put in place annual model API capacity for an AI product, with usage forecasts, provider fit, quotas, cost controls and the integration work that makes the capacity usable.
Where this starts
The operating problem
The engagement focuses on the surrounding system: boundaries, evidence, permissions, exceptions, people and the decisions the implementation must support.
Capacity is estimated from user counts without modelling requests, tokens or retries
Model, prompt, retrieval and caching choices change the cost profile materially
Rate limits, project quotas and budget alerts do not match the product boundary
Procurement is separated from architecture, evaluation and release planning
What leaves the engagement
Ways to start
1–2 weeks
Translate product traffic and AI workflows into a defensible model-usage and cost range.
2–4 weeks
Set up the agreed capacity, application integration, project controls, monitoring and handover.
12 months
Combine agreed inference capacity with usage review, quota management and optimisation through the term.
Inside the work
Representative engagements show how this capability connects to the product, workflow, controls, and people around it.
Service questions
No. The useful scope connects commercial capacity to the workload, provider configuration, application integration, quotas, monitoring and cost controls. Capacity without those decisions is difficult to operate.
The provider follows the use case, data constraints, model quality, regional availability and commercial terms. The final scope names the provider and what its credits, quotas and support actually include.
Start with context
Describe the current system, the decision ahead and the constraint that is making progress difficult.