Billing granularity and what it covers
Baseten bills dedicated deployments per minute and charges for deployment time as well as request time, though it does not bill for idle time between requests. Modal bills per second of compute. RunPod bills by GPU hour with separate serverless and on-demand pod rates. Replicate bills most public models per second of run time and private models for all the time an instance is online, including setup and idle. A team whose requests last a few seconds should weigh whether per-minute rounding adds meaningful overhead, while a team running sustained inference should compare effective hourly rates directly.