What It Does
Together AI provides fast, affordable inference and training for open-source language models. Their optimized infrastructure serves models like Llama, Mistral, and Qwen at speeds and prices that undercut most competitors.
Key Features
- Fast inference — optimized serving with custom kernels and hardware
- Leading open models — Llama, Mistral, Qwen, DeepSeek immediately available
- Fine-tuning — customize models on your data
- Custom models — deploy your own models on Together infrastructure
- Embeddings — text embedding models for RAG applications
- Function calling — structured output and tool use
- Competitive pricing — aggressive cost optimization
Pricing Breakdown
| Model Class | Input | Output |
|---|---|---|
| Small (8B) | $0.10/M | $0.10/M |
| Medium (70B) | $0.54/M | $0.54/M |
| Large (405B) | $2.56/M | $2.56/M |
Who It’s For
Developers building applications on open-source models who need fast inference at low cost, startups that want GPT-4-class capability without GPT-4 pricing, and teams fine-tuning models for specific use cases.
Competitive Position
Together AI competes on speed and price for open-source model inference. Their custom serving infrastructure often delivers the fastest inference for popular models. Competes with Fireworks AI (similar positioning), Groq (hardware advantage), and cloud providers (more expensive). The research background (founding team includes academics) informs their optimization work.