Per-token. No minimums. No surprises.
USD at launch. Local currency display (NGN, KES, ZAR, EGP, GHS, EUR) ships with the local-rails wave. Pay only for what you use — most teams spend less than $50/month.
Embeddings
Where the margin lives. Multi-lingual, multi-granularity.Rerankers
Cross-encoder scoring for retrieval pipelines.Chat & reasoning
Small, sharp, frontier-class. We don't host 405B-vanity.Speech
Per-minute audio in.All prices in USD per 1M tokens unless noted. EU egress free between Tomoul regions; standard pass-through to other clouds.
Free
- 10 req/s rate limit
- All public models
- Community support
- Shared infrastructure
Pay-as-you-go
- 500 req/s rate limit
- Auto-top-up + invoicing
- Email + Discord support
- Multi-region routing
- Local-currency billing
Enterprise
- Dedicated capacity
- Custom regions (Lagos, Nairobi, Cairo)
- SOC 2 + DPA + sub-processor list
- SLA · 99.95% uptime
- Slack channel
Pricing FAQ
How does per-token billing work?
You pay only for the input and output tokens you actually send. We meter per request and aggregate hourly into your account balance. No minimums, no idle GPU charges, no commit.
What about local currency?
USD at launch. NGN, KES, ZAR, EGP, GHS, and EUR billing display ships with the local-rails wave (Flutterwave, Paystack, M-Pesa) alongside the Lagos POP in H2 2026. You'll be able to pay your local merchant of record, not Stripe.
Is there a free tier?
$5 free credit on signup, no card required. After that you top up with Stripe or USDC. Credits never expire.
Do you train on my data?
No. Never. By default we don't log prompts or completions, and we have no data-sharing arrangements with model providers. See /trust for residency and processor commitments.
What's the cheapest way to use Tomoul?
If your workload is privacy-sensitive or sporadic, run `tomoul serve` locally. The same CLI points at our cloud with `--cloud` when you need scale — same models, same API.