| Example | Description |
|---|---|
| Aimlapi | AIML API examples for basic runs, multimodal input, memory, retries, structured output, and tool use. |
| Anthropic | Claude examples for multimodal input, context management, caching, knowledge, memory, thinking, structured output, server tools, and skills. |
| AWS | Run Claude and Amazon Nova models on AWS Bedrock. |
| Azure | Run Claude and open-source models on Azure AI Foundry and OpenAI endpoints. |
| Cerebras | Cerebras examples: basic runs, storage, knowledge, structured output, retries, and tool use. |
| Cerebras OpenAI | Run Cerebras models through the OpenAI-compatible endpoint: streaming, tools, structured output, storage, and knowledge. |
| Clients | Configure Agno’s default sync httpx.Client with headers, logging, request IDs, timeouts, and error tracking. |
| Cohere | Run Cohere Command A and Aya Vision models with tools, knowledge, memory, retries, and structured output. |
| Cometapi | Run GPT, Claude, Gemini, DeepSeek, and Qwen models through CometAPI’s OpenAI-compatible gateway. |
| Dashscope | Browse DashScope model examples with Qwen models, image analysis, knowledge tools, and retry patterns. |
| DeepInfra | DeepInfra examples: basic agent runs, JSON output, tool use, and retries. |
| DeepSeek | Run DeepSeek models with reasoning, thinking mode, structured output, retries, and tool use. |
| Fireworks | Run Fireworks models with streaming, structured output, web search, and retry configuration. |
| Use Gemini for audio, video, image, PDF, grounding, file search, and thinking-budget examples. | |
| Groq | Groq examples for agents and teams, multimodal input, knowledge, reasoning, research, transcription, translation, structured output, and tools. |
| Hugging Face | Hugging Face examples for basic and streaming runs, essay generation, retries, and web-search tool use. |
| IBM | IBM watsonx examples for model retries, storage, knowledge, structured output, and tools. |
| Internlm | Run InternLM models with basic responses, tools, knowledge, storage, retries, and structured output. |
| LangDB | Run LangDB models with basic responses, tools, retries, and structured output. |
| LiteLLM | Run agents through the LiteLLM gateway with tools, knowledge, structured output, and audio, image, and PDF input. |
| LiteLLM OpenAI | Examples for LiteLLM with OpenAI-compatible models. |
| Llama Cpp | Run agents against a local llama.cpp server serving ggml-org/gpt-oss-20b-GGUF at http://127.0.0.1:8080/v1. |
| Lmstudio | LM Studio examples for local models, images, knowledge, memory, storage, retries, structured output, and tools. |
| Meta | Llama and Llama OpenAI examples covering tool use, knowledge, memory, metrics, storage, and retries. |
| Mistral | Run Mistral models with image input, memory, structured output, retries, and tool use. |
| Moonshot | Moonshot Kimi K2 agent examples: basic sync/streaming responses and web-search tool use. |
| N1N | N1N gateway examples: running OpenAI models via N1N with basic streaming and web-search tool calls. |
| Nebius | Nebius model examples: basic runs, Postgres sessions, PgVector knowledge, retries, structured output, and tool use. |
| Neosantara | Neosantara examples covering basic runs, structured output, and web-search tool use. |
| Nexus | Nexus examples covering basic runs, retry configuration, and tool use. |
| NVIDIA | NVIDIA API examples: basic runs, retry configuration, and tool use. |
| Ollama | Ollama Chat and Responses API examples for local and cloud models, knowledge, memory, reasoning, structured output, and tools. |
| OpenAI | OpenAI Chat and Responses API examples for multimodal input, tools, reasoning, structured output, storage, and streaming. |
| OpenRouter | OpenRouter Chat and Responses API examples for model routing, retries, structured output, and tools. |
| Perplexity | Index of Perplexity sonar-pro agent examples: basic runs, knowledge, memory, retries, structured output, and web search. |
| Portkey | Index of Agno examples routing agents through the Portkey AI gateway: basic runs, retries, structured output, and tool use. |
| Requesty | Requesty AI is an LLM gateway with AI governance. See their website for more information. |
| Sambanova | SambaNova model examples: basic sync/stream/async runs and retry configuration. |
| Siliconflow | Examples for SiliconFlow model integration. |
| Together | Run Together models with streaming, image input, reasoning, structured output, web search, and retry configuration. |
| Vercel | Index of Agno examples running on Vercel’s v0 model: basic runs, images, knowledge, retries, and web search. |
| Vertex AI | Vertex AI examples for Claude models, retries, multimodal input, knowledge, memory, caching, structured output, and tools. |
| vLLM | vLLM is a fast and easy-to-use library for running LLM models locally. |
| xAI | xAI model examples for building agents with Grok, including vision, web search, and financial analysis. |
| Cloudflare | Cloudflare AI Gateway model examples. |
| Inception | Inception Labs Mercury model examples. |
| MiniMax | MiniMax M3 agent examples: basic runs, web search tool use, and JSON-mode structured output. |
| Xiaomi MiMo | Xiaomi MiMo model examples. |
Models
Models
Examples for all supported LLM providers in Agno.