ChatCerebras chat models. For detailed documentation of all ChatCerebras features and configurations head to the API reference.
At Cerebras, we’ve developed the world’s largest and fastest AI processor, the Wafer-Scale Engine-3 (WSE-3). The Cerebras CS-3 system, powered by the WSE-3, represents a new class of AI supercomputer that sets the standard for generative AI training and inference with unparalleled performance and scalability.
With Cerebras as your inference provider, you can:
- Achieve unprecedented speed for AI inference workloads
- Build commercially with high throughput
- Effortlessly scale your AI workloads with our seamless clustering technology
Overview
Integration details
Model features
Setup
Credentials
Get an API Key from cloud.cerebras.ai and add it to your environment variables:Installation
The LangChain Cerebras integration lives in thelangchain-cerebras package:
Instantiation
Now we can instantiate our model object and generate chat completions:Invocation
Streaming
Async
Async streaming
API reference
For detailed documentation of allChatCerebras features and configurations head to the API reference: API reference
Connect these docs to Claude, VSCode, and more via MCP for real-time answers.

