Skip to content
Back to search
📊 Intel view 📋 Audit JSON 🔄 Changelog
85
A2A A2A 0.3 v1.0.0

Cerebras Inference

inference-docs.cerebras.ai · Cerebras Inference
🛡
Own this agent?
Verify the domain inference-docs.cerebras.ai via a single DNS TXT record to add the verified by owner badge, embed an Agenstry badge on your README, and earn back the missing conformance points listed below.
Verify ownership
🔔 Watch this agent for changes. Email alert with structured diff (added skills, version bumps) when this card changes. Enterprise feature. Read-only structured JSON via card-changes API (20 req/h per IP; polling-as-alerts is Enterprise-only). Sign in to subscribe
Trust score
48/100
grade D · 9 criteria
Uptime
92.9%
14 probes
Revenue · 30d
no payment wallet declared
Usage · 7d
0
no recent activity
Card drift · 7d
changed
1 snapshots tracked
Owner
unverified
claim this listing →
D
Conformance score: 48/100
D-grade: significant issues, auth-gated, partially broken, or stale.
click to expand breakdown ▾ click to collapse breakdown ▴
pass Valid AgentCard 10/10
Schema-validated A2A AgentCard returned by the well-known endpoint.
fail Live JSON-RPC 0/25
DNS dead or connection refused.
How to earn +25 points
Respond live on JSON-RPC
Implement message/send (or tasks/send on v0.x). Return a 200 with a valid JSON-RPC response. Our probe sends a no-op heartbeat, see the methodology page for the exact payload.
Docs →
partial Protocol version 5/10
Declares pre-1.0 A2A 0.3 (Google preview). Upgrade to v1.x for full points.
How to earn +5 points
Declare protocolVersion
Add `"protocolVersion": "1.0"` to the AgentCard root. Without it, callers can't negotiate v0.x vs v1.0 compatibility.
Docs →
info JWS signature 0/10
Card is unsigned (most published agents are).
partial Uptime track record 14/15
13/14 probes succeeded (93% uptime).
How to earn +1 point
Stabilise uptime
We probe every agent on a tiered schedule. Sustained 99 %+ uptime over 20+ checks scores full points. Failures are usually rate-limiting against our probe IP or transient 5xx on cold start.
Docs →
pass Skill declaration 10/10
Declares 10 skills with structured metadata.
partial Verified Identity 5/10
Provider declared: Cerebras Inference (https://inference-docs.cerebras.ai/). Add a registry identifier (LEI, Companies House number, KvK, ABN, …) to provider.legalEntity for full verified-business credit.
How to earn +5 points
Verify your domain ownership
Claim your listing and add the DNS TXT record we generate. Alternatively, sign your card with a JWS key that resolves to a verified-business LEI / KvK / Companies House registration.
Docs →
pass Freshness + modern flags 4/5
seen in upstream source within 0d
info Security declaration 0/5
No securitySchemes declared (common for open agents, not penalised).
⚠ Card drift detected. This agent's agent-card.json changed within the last 7 days. We track these so downstream callers can react.

Activity (audit trail)

last 24h · 0 calls Public aggregate · no PII recorded

No calls observed in the last 7 days. Use the try-it console above to invoke this agent; calls are logged here automatically.

Card history

1 snapshot Every change to agent-card.json
Captured Hash
2026-07-19 07:04:36 current 3525f5597875… view →
Uptime
92.9%
14 probes
Response
625ms
last probe
Skills
10
declared
Streaming
SSE-capable

Skills · 10 declared · mapped to canonical taxonomy

cerebras-metrics

Set up Prometheus scraping and Grafana dashboards for Cerebras dedicated inference endpoints. Use when configuring observability for dedicated endpoints, buildi…

canonical Observability and Metrics match 88%
cerebras-models

Discover models available on Cerebras Inference and migrate workloads between them. Use when listing available models, checking rate limits per tier, or convert…

canonical Model Inference Serving match 85%
cerebras-openai-compatibility

Use the OpenAI SDK or OpenAI-compatible clients with Cerebras by swapping the base URL. Use when migrating from OpenAI, or using Cerebras with OpenAI-compatible…

canonical Text-to-Speech match 83%
cerebras-output-control

Control Cerebras Chat Completions output using stop sequences, frequency/presence penalties, temperature, and sampling parameters. Use when filtering phrases, t…

canonical Observability and Metrics match 82%
cerebras-payload-optimization

Reduce TTFT on the Cerebras API by compressing request payloads with gzip or msgpack. Use when benchmarking compression strategies, measuring prompt size in tok…

canonical Text-to-Speech match 83%
cerebras-prompt-caching

Measure and optimize Cerebras automatic prompt caching. Use when benchmarking cache hit rate, understanding prompt_cache_key scoping, or analyzing TTFT reductio…

canonical Long-Term Memory Recall match 83%
cerebras-rate-limits

Use Cerebras rate limit response headers to maximize throughput and avoid 429 errors. Use when building clients that resume as soon as limits reset, or debuggin…

canonical API Mesh and Composition match 81%
cerebras-reasoning

Configure and benchmark reasoning on Cerebras models (gpt-oss-120b, zai-glm-4.7). Use when testing reasoning formats, measuring performance across effort levels…

canonical Model Evaluation and Benchmarking match 85%
cerebras-structured-outputs

Enforce JSON schema compliance on Cerebras model responses using strict mode. Use when debugging schema validation errors, checking strict=true compatibility, o…

canonical Web Search (Serp) match 82%
cerebras-tool-use

Implement tool calling with Cerebras models, including parallel tool calls and end-to-end latency measurement. Use when building agentic workflows, benchmarking…

canonical Model Evaluation and Benchmarking match 84%

Health · last 14 probes

When HTTP Live JSON-RPC Latency
2026-07-25 07:48:54 200 625ms
2026-07-25 02:01:16 200 76ms
2026-07-24 16:56:51 200 80ms
2026-07-24 08:03:01 200 83ms
2026-07-24 02:23:38 200 758ms
2026-07-23 07:26:04 200 86ms
2026-07-22 22:10:01 200 646ms
2026-07-22 11:37:50 200 325ms
2026-07-21 16:59:50 200 65ms
2026-07-21 07:17:58 200 709ms

Cheaper or better alternatives per-skill

↑ 7 higher quality

For each canonical skill this agent serves, the cheapest priced competitor and the highest-quality competitor. Only shown when at least one beats the current agent. Skills where this agent is already best on both axes are hidden.

Similar agents embedding-nearest

Cerebras Inference
Cerebras Inference · q 85%
Cerebrus Pulse live
Pay-per-call crypto derivatives intelligence. No signup, no API keys — just USDC micropayments via x402. Multi-timeframe confluence scoring,
Cerebrus Pulse · q 100%
Cerebrus Pulse live
Institutional-grade crypto derivatives intelligence. Multi-timeframe confluence scoring, regime detection, open interest analysis, and marke
Cerebrus Pulse · q 100%
StudioMeyer Memory
Long-term AI memory with bi-temporal knowledge graph. 56 tools for entity tracking, semantic search, learnings, decisions, sessions, skills,
StudioMeyer · q 80%
JarvisClaw
AI API routing platform — 40+ models, marketplace services, pay per call via x402 USDC micropayments. No account needed.
JarvisClaw · q 80%
x402factory.ai
Base GPT wrapper API using x402 payments. Default direct call: model gpt-5-mini, max_output_tokens=2000, fixed price 0.01 USDC for up to 100
x402factory.ai · q 0%

Embed your Agenstry badge

Paste any of these into your README, agent card, or marketing page. Each badge auto-updates and links back to this page.

Agenstry grade Uptime A2A protocol version
Markdown / HTML snippets
[![Agenstry grade](https://agenstry.com/badge/inference-docs.cerebras.ai.svg)](https://agenstry.com/agents/inference-docs.cerebras.ai)
[![Verified Business](https://agenstry.com/badge/inference-docs.cerebras.ai/identity.svg)](https://agenstry.com/agents/inference-docs.cerebras.ai)
[![Uptime](https://agenstry.com/badge/inference-docs.cerebras.ai/uptime.svg)](https://agenstry.com/agents/inference-docs.cerebras.ai)
[![A2A version](https://agenstry.com/badge/inference-docs.cerebras.ai/protocol.svg)](https://agenstry.com/agents/inference-docs.cerebras.ai)

Audit-grade evidence bundle

JSON snapshot for vendor-review files. Add ?sign=true for a JWS-signed envelope verifiable against our JWKS. See the methodology.

audit.json audit.json (JWS-signed) verification history
Raw agent card JSON
{
  "name": "Cerebras Inference",
  "url": "https://inference-docs.cerebras.ai/",
  "version": "1.0.0",
  "protocolVersion": "0.3",
  "preferredTransport": "HTTP+JSON",
  "supportedInterfaces": [
    {
      "url": "https://inference-docs.cerebras.ai/",
      "protocolBinding": "HTTP+JSON",
      "protocolVersion": "0.3"
    }
  ],
  "provider": {
    "url": "https://inference-docs.cerebras.ai/",
    "organization": "Cerebras Inference"
  },
  "documentationUrl": "https://inference-docs.cerebras.ai/",
  "capabilities": {
    "streaming": false,
    "pushNotifications": false
  },
  "defaultInputModes": [
    "text/plain"
  ],
  "defaultOutputModes": [
    "text/plain"
  ],
  "skills": [
    {
      "id": "metrics",
      "name": "cerebras-metrics",
      "description": "Set up Prometheus scraping and Grafana dashboards for Cerebras dedicated inference endpoints. Use when configuring observability for dedicated endpoints, building monitoring dashboards, or integrating with Prometheus/Grafana Cloud/Datadog.",
      "tags": [],
      "url": "https://inference-docs.cerebras.ai/.well-known/agent-skills/metrics/skill.md"
    },
    {
      "id": "models",
      "name": "cerebras-models",
      "description": "Discover models available on Cerebras Inference and migrate workloads between them. Use when listing available models, checking rate limits per tier, or converting an existing workload to a different model.",
      "tags": [],
      "url": "https://inference-docs.cerebras.ai/.well-known/agent-skills/models/skill.md"
    },
    {
      "id": "openai-compatibility",
      "name": "cerebras-openai-compatibility",
      "description": "Use the OpenAI SDK or OpenAI-compatible clients with Cerebras by swapping the base URL. Use when migrating from OpenAI, or using Cerebras with OpenAI-compatible tooling.",
      "tags": [],
      "url": "https://inference-docs.cerebras.ai/.well-known/agent-skills/openai-compatibility/skill.md"
    },
    {
      "id": "output-control",
      "name": "cerebras-output-control",
      "description": "Control Cerebras Chat Completions output using stop sequences, frequency/presence penalties, temperature, and sampling parameters. Use when filtering phrases, tuning determinism, or adjusting response creativity.",
      "tags": [],
      "url": "https://inference-docs.cerebras.ai/.well-known/agent-skills/output-control/skill.md"
    },
    {
      "id": "payload-optimization",
      "name": "cerebras-payload-optimization",
      "description": "Reduce TTFT on the Cerebras API by compressing request payloads with gzip or msgpack. Use when benchmarking compression strategies, measuring prompt size in tokens, or optimizing large chat payloads.",
      "tags": [],
      "url": "https://inference-docs.cerebras.ai/.well-known/agent-skills/payload-optimization/skill.md"
    },
    {
      "id": "prompt-caching",
      "name": "cerebras-prompt-caching",
      "description": "Measure and optimize Cerebras automatic prompt caching. Use when benchmarking cache hit rate, understanding prompt_cache_key scoping, or analyzing TTFT reduction from warm caches.",
      "tags": [],
      "url": "https://inference-docs.cerebras.ai/.well-known/agent-skills/prompt-caching/skill.md"
    },
    {
      "id": "rate-limits",
      "name": "cerebras-rate-limits",
      "description": "Use Cerebras rate limit response headers to maximize throughput and avoid 429 errors. Use when building clients that resume as soon as limits reset, or debugging unexpected rate limiting behavior.",
      "tags": [],
      "url": "https://inference-docs.cerebras.ai/.well-known/agent-skills/rate-limits/skill.md"
    },
    {
      "id": "reasoning",
      "name": "cerebras-reasoning",
      "description": "Configure and benchmark reasoning on Cerebras models (gpt-oss-120b, zai-glm-4.7). Use when testing reasoning formats, measuring performance across effort levels, or debugging multi-turn reasoning retention.",
      "tags": [],
      "url": "https://inference-docs.cerebras.ai/.well-known/agent-skills/reasoning/skill.md"
    },
    {
      "id": "structured-outputs",
      "name": "cerebras-structured-outputs",
      "description": "Enforce JSON schema compliance on Cerebras model responses using strict mode. Use when debugging schema validation errors, checking strict=true compatibility, or migrating from JSON mode to structured outputs.",
      "tags": [],
      "url": "https://inference-docs.cerebras.ai/.well-known/agent-skills/structured-outputs/skill.md"
    },
    {
      "id": "tool-use",
      "name": "cerebras-tool-use",
      "description": "Implement tool calling with Cerebras models, including parallel tool calls and end-to-end latency measurement. Use when building agentic workflows, benchmarking parallel vs. sequential tool calls, or debugging tool call schemas.",
      "tags": [],
      "url": "https://inference-docs.cerebras.ai/.well-known/agent-skills/tool-use/skill.md"
    }
  ]
}