Command (Cohere)
Overview
Cohere is an enterprise-focused AI company. Founded in 2019 in Toronto, Cohere has carved a unique position as an LLM provider specializing in enterprise RAG, retrieval, and agents rather than consumer chat. Released in 2025, Command A delivers 256K context, 23 languages, and runs on just 2 A100/H100 GPUs.
Latest Model: Command A (2025)
Cohere's most performant model, specialized for RAG, tool use, agents, and multilingual use cases.
Performance & Efficiency
| Metric | Value |
|---|---|
| GPU Requirement | Just 2 A100 or H100 (vs up to 32 for others) |
| Speed | 156 tok/sec (1.75x faster than GPT-4o) |
| Context Length | 256,000 tokens (~600 pages) |
| Languages | 23 languages (including improved Arabic dialects) |
Key Features
1. Efficient GPU Requirements
Runs on just 2 A100/H100 GPUs — far fewer than the 32 required by some competitors. Dramatically lowers private deployment costs.
2. Enterprise RAG Specialization
Optimized for high-accuracy retrieval-augmented generation. Produces verifiable answers with citations.
3. Agentic Tool Use
Supports complex workflows with agent-style tool use.
4. North AI Platform
Cohere's enterprise agent platform for secure in-house automation.
5. Model Vault (Launched September 2025)
Dedicated inference platform. Deploy Command, Rerank, and Embed in isolated VPCs or on-premises — sensitive data never leaves your organization.
6. Command A Vision
Cohere's first vision-capable model. Enterprise-grade image understanding.
Specialized APIs (Key to RAG)
Rerank API
Specialized for reranking search results. Significantly improves RAG pipeline accuracy.
Embed API (Multilingual)
23-language multilingual embedding model. Cohere is recognized as industry-leading in multilingual RAG.
Pricing
| Item | Price |
|---|---|
| Command A API | Enterprise pricing (contact) |
| Command R+ API | $2.50 / $10 per 1M tokens |
| Rerank / Embed API | Pay-per-use |
| Private deployment | Up to 50% cheaper than API |
Strengths
- ✅ High-accuracy enterprise RAG
- ✅ 23-language multilingual embeddings
- ✅ Runs on just 2 GPUs (highly efficient)
- ✅ Model Vault for on-prem / VPC deployment
- ✅ North AI platform for agentic automation
- ✅ 1.75x faster than GPT-4o
Weaknesses
- ❌ No consumer product (no ChatGPT/Claude.ai-like app)
- ❌ Lower brand awareness than OpenAI/Anthropic
- ❌ No image/video generation
- ❌ No consumer free tier
Official Resources
- Website: https://cohere.com
- API Docs: https://docs.cohere.com
- Command A: https://docs.cohere.com/docs/command-a
- North AI: https://cohere.com/north