Cerebras API
Cerebras 接口
Fastest LLM inference. Llama 3.1 405B at 969 tokens/second with OpenAI-compatible API.
最快的 LLM 推理。Llama 3.1 405B 达 969 tokens/秒,兼容 OpenAI API。
Website
Documentation
Pricing
free tier (1M tokens/day) + pay-as-you-go
Authentication
- API Key
Popularity
54/100
Endpoints
https://api.cerebras.ai/v1/chat/completions
SDKs
PythonOpenAI SDK compatible
Tags
llmfast-inferenceopen-source
Last updated: 2026-08-01
Similar APIs
OpenAI API
★100Access GPT-4o, GPT-4, GPT-3.5, embeddings, DALL-E, Whisper, and TTS models via REST API.
GitHub REST API
★96Manage repos, issues, PRs, actions, and users. The backbone of open-source automation.
Hugging Face Hub API
★92Access 500K+ models, datasets, and Spaces. Inference API for instant model deployment.
LangChain Platform
★91Framework and platform for building LLM applications. LangSmith for observability, LangServe for deployment.
Something wrong or missing? Open an issue on GitHub →
Pricing Comparison
| API | Pricing | Auth | Popularity |
|---|---|---|---|
| OpenAI API | pay-as-you-go (per token) | API Key (Bearer Token) | 100 |
| Anthropic Claude API | pay-as-you-go (per token) | API Key (x-api-key header) | 87 |
| Google Gemini API | free tier available, pay-as-you-go | API Key | 85 |
| OpenAI Assistants API | pay-as-you-go (per token + per tool use) | API Key (Bearer Token) | 84 |
Discussion
Questions, feedback, or suggestions? Start a discussion.
Discuss on GitHub