Appearance
Curated Resource List
The goal of this list is not to be exhaustive—it is to make sure every entry is worth your time. There are only three inclusion criteria: is the content first-hand (official docs, the author's own blog, the original repo), has it been tested in real engineering practice (production case studies, runnable code), and is it still maintained (substantive updates after 2025). An entry must meet at least two of the three to make the cut.
The list is organized as a learning path: start with official docs and the classic posts to build judgment, then pick a framework and get your hands dirty, then calibrate your understanding with benchmarks, and finally rely on information sources to stay current. If you are not sure what to learn first, work through the learning paths page and come back to follow this list step by step.
How to use this list
Information in the Agent field has a half-life of about 6 months. This list carries a dataAsOf stamp (2026-08); framework star counts and course availability may have changed, so check the original pages before quoting specific numbers. More important than bookmarking links: make a fixed habit of reading 2-3 first-hand sources every week—it is far more efficient than skimming second-hand news.
1. Official Documentation and Cookbooks
Official documentation is the only source in this field that will not mislead you—vendors are accountable for their own APIs. Third-party tutorials frequently copy from outdated API versions; official docs do not. The three vendors each have a distinct style, so pick what fits: Anthropic's docs lean toward design philosophy and best practices, OpenAI's toward guides and sample code, and Google's toward a structured reference manual.
Anthropic
- Claude official documentation — the authoritative reference for the Agent SDK, tool calling, and MCP integration. Note that the domain has moved from docs.anthropic.com to docs.claude.com, so it is time to update old bookmarks. For anyone developing against Claude-family models.
- Claude Cookbooks — an officially maintained collection of copy-and-run code snippets covering RAG, tool use, and Agent patterns, all runnable as-is. Ideal for hands-on learners who learn by adapting working code.
- Claude Code documentation — even if you never plan to use Claude Code, its docs are the best observational sample of what configuration surface a mature Coding Agent exposes. It reads even better paired with this site's Claude Code case study.
OpenAI
- OpenAI Agents SDK documentation — official docs for the lightweight Agent framework, with clean definitions of handoffs, guardrails, and tracing. See this site's OpenAI Agents SDK page for details.
- OpenAI Cookbook — the official examples repo. The
examples/agents_sdk/directory contains complete notebooks spanning single-Agent setups to multi-Agent collaboration, including a side-by-side guide for migrating from the Claude Agent SDK—worth reading when you are weighing the two SDKs against each other. - A Practical Guide to Building Agents — OpenAI's enterprise guide to building Agents, published in 2025 (PDF). It breaks an Agent down into three elements—model, tools, and instructions—and offers criteria for when to evolve from a single Agent to multi-Agent designs. Written for technical decision-makers, but engineers get plenty out of it too.
A taste of the official SDK's minimalist style (adapted from the quickstart in the official docs; treat the docs as authoritative for details):
python
from agents import Agent, Runner
# Define an Agent: name, instructions, and available tools
agent = Agent(
name="assistant",
instructions="You are a rigorous assistant: state the conclusion first, then the evidence.",
)
# Runner drives the Agent loop; run_sync is the synchronous entry point
result = Runner.run_sync(agent, "Summarize this week's LLM news for me")
print(result.final_output)Ten lines or fewer and you have a working Agent—which is why the post-2025 debate shifted from "should you use a framework" to "should you use a heavyweight framework." Lightweight SDKs have already compressed boilerplate to nearly zero.
Google
- Agent Development Kit (ADK) documentation — Google's Agent framework, open sourced in 2025. The docs are top-tier by big-vendor standards, with a separate repository for each language implementation (Python/Java/Go/TS).
- google/adk-python — the main ADK repo. Its samples directory deserves more attention than the docs themselves.
Protocol Layer
- Model Context Protocol site and specification — home of the first-party MCP spec. The specification directory keeps dated version snapshots (such as the 2025-06-18 edition). Read the spec, not blog posts, before writing an MCP server. This site's Tools and MCP page has a beginner-friendly walkthrough.
- modelcontextprotocol GitHub organization — where the official SDKs and reference server implementations live.
2. Books and Courses
Books
Few books devoted specifically to Agents hold up under scrutiny, so for now only one makes the cut as a foundation:
AI Engineering by Chip Huyen (O'Reilly, 2025) — strictly speaking this is not an Agent book but a panoramic engineering guide to building applications on foundation models. Even so, its chapters on evals, RAG, fine-tuning, and inference costs are precisely the foundations of Agent engineering. The companion repo chiphuyen/aie-book carries the author's curated resources. Best for engineers who want a systems perspective rather than a tour of frameworks. A Chinese translation is also in print—check the edition details on the bookstore page before buying.
Suggested reading order: if you are short on time, read the evals and RAG chapters first—the twin cornerstones of Agent engineering. Skim the prompt engineering chapter, and leave fine-tuning until your business actually hits a bottleneck.
Why no more books
Agent engineering only began converging on a shared consensus at the end of 2024 (Anthropic's Building effective agents was published in December 2024), and the publishing cycle of a printed book cannot keep up with that pace. For now, blogs plus official docs plus papers are better vehicles; books are for shoring up foundations, not for chasing the frontier.
Free Courses
- Hugging Face Agents Course — a free, open source Agent course with a certificate, built around smolagents to teach the Code Agent paradigm. The companion site lives under huggingface.co/learn, and a community-maintained Chinese translation is available. Ideal for a systematic start from zero.
- Microsoft AI Agents for Beginners — Microsoft's beginner course, released in February 2025 and since expanded from 10 to 18 lessons. It covers design patterns such as tool use, Agentic RAG, planning, and multi-Agent systems, with an official Simplified Chinese version. Roughly 50k+ GitHub stars; a good fit for beginners with coding experience who can follow the lesson plan.
- DeepLearning.AI: AI Agents in LangGraph — a short course taught personally by LangChain founder Harrison Chase (about an hour of material), moving from a hand-written ReAct loop to state and persistence in LangGraph. Free, and a good first contact with LangGraph.
- The DeepLearning.AI short-course series as a whole deserves attention: the same platform hosts other Agent-related short courses such as Functions, Tools and Agents with LangChain. They are small, free, and handy for filling gaps in spare moments.
- Google's 5-Day AI Agents Intensive, run with Kaggle — a live, bootcamp-style program centered on ADK, with course materials archived on Kaggle for self-paced study. A good fit if you want to stay on Google's stack.
About certificates
Certificates from these courses carry little weight in the job market—what interviewers want to see is what you built with the patterns the courses teach. Rework, extend, and deploy each course's capstone project and put it in your portfolio; that is worth far more than the certificate itself. See this site's resume analysis.
3. Classic Engineering Blog Posts
If papers tell you what is possible, engineering blogs tell you what actually works in production. The pieces below are the most-cited writings from 2024-2026 whose value has stood the test of time, ordered by how essential they are.
There is a method for reading this kind of article: do not just record the conclusions—record the boundaries within which they hold. For example, "multi-Agent beats single-Agent" only holds when tasks can be parallelized and the information exceeds a single context window; Anthropic says plainly in the same piece that most coding tasks are actually a poor fit for multi-Agent. Extracting these boundary conditions pays off in interviews and architecture reviews alike.
| Article | Source and Date | Value in One Line |
|---|---|---|
| Building effective agents | Anthropic, 2024-12 | The most-cited engineering article in this field, bar none. Its core claim—"most scenarios do not need a complex framework; simple composable patterns suffice"—still holds today |
| Context Engineering for AI Agents: Lessons from Building Manus | Manus (Ji Yichao), 2025-07 | From KV-cache hit rates to "the file system as context," every conclusion here was paid for in production scars. Cited by this site's context engineering page and the Manus case study |
| How we built our multi-agent research system | Anthropic, 2025-06 | A retrospective on the multi-Agent architecture behind Claude Research, with hard conclusions such as "token consumption explains 80% of the performance variance." Read alongside multi-agent architecture |
| Effective context engineering for AI agents | Anthropic, 2025 | Turns context engineering from a slogan into a methodology: what belongs in the context and what should stay out |
| Context Engineering for Agents | LangChain, 2025 | Organizes context strategies into four types—write, select, compress, isolate—one of the most usable taxonomies around |
| Writing effective tools for agents | Anthropic, 2025 | A hands-on guide to tool design—writing tools for Agents and writing APIs for humans are two different crafts |
| Building agents with the Claude Agent SDK | Anthropic, 2025-09 | The official guide to building with the same SDK that powers Claude Code. Its guiding idea—"hand it a harness with a working loop, not a pile of rules"—is highly representative. Pairs with this site's Claude Agent SDK page |
| Demystifying evals for AI agents | Anthropic, 2026-01 | A systematic guide to Agent evaluation: the three-way grader taxonomy, the difference between pass@k and pass^k, and how to evaluate each Agent type (coding / conversational / research / computer use). Read alongside this site's evaluation framework |
| 12-Factor Agents | HumanLayer (Dex Horthy), 2025 | A production principles list with an anti-framework stance; entries like "own your context window" and "small, focused agents" have become industry mantras. Published as a GitHub repo—worth printing and pinning to the wall |
| Claude Code best practices | Anthropic, 2025-04 | Nominally about CLI usage, but in substance a mental model for collaborating effectively with a Coding Agent |
| Unrolling the Codex agent loop | OpenAI, 2025 | The Codex CLI author personally dissects the Agent Loop implementation: how prompt prefixes align with the cache, and how to compact context once it exceeds the threshold. Read alongside this site's Agent Loop page |
| Effective harnesses for long-running agents | Anthropic, 2025-11 | How to make long-horizon Agents that run for hours execute reliably: harness design, checkpoint recovery, and cross-session memory. Long-horizon tasks are the main battleground of 2026 |
Two more blogs are worth subscribing to rather than reading piecemeal:
- Anthropic Engineering — infrequent updates, but every post delivers. Half of the classics above came from here.
- The LangChain blog — inherently promotional (it is their own framework, after all), but its "agent concept clarifications" and "production case interview" series are consistently good. Just keep facts and marketing separate as you read.
4. Open Source Frameworks and Repos
Star counts are order-of-magnitude figures from mid-2026, taken from public statistics pages. GitHub stars change constantly, so treat the repository homepages as authoritative. For the selection logic see this site's framework selection overview; this section is navigation only.
First, a take: the 2026 framework landscape has clearly stratified. The bottom layer is the vendors' official SDKs (lightweight, close to the API), the middle layer is orchestration frameworks (LangGraph, CrewAI—the kind that manage state and flow), and the top layer is low-code platforms. Most production teams actually choose either "official SDK plus a little homegrown orchestration" or "the full LangGraph stack," and frameworks straddling the middle are being squeezed. Before picking a framework, work out which layer your problem actually lives in.
Python Orchestration Frameworks
| Repo | Star Count | Positioning and Fit |
|---|---|---|
| langchain-ai/langchain | 130k+ | The LLM application framework with the largest ecosystem and the most integrations; for when you need to quickly assemble ready-made components |
| langchain-ai/langgraph | Tens of thousands (roughly 30-50k) | Graph-based, stateful Agent orchestration and a common default for production Agents; for teams that need fine-grained control over flow |
| openai/openai-agents-python | 10k+ | A lightweight SDK built around the handoff/guardrail/tracing trio; for getting started fast inside the OpenAI ecosystem |
| crewAIInc/crewai | 50k+ | Role-playing multi-Agent; prototypes in a few dozen lines of code. For business process automation and demos. See this site's CrewAI page |
| microsoft/autogen | 50k+ | Conversational multi-Agent framework. Note: the official repo is now in maintenance mode—Microsoft has shifted its focus to Microsoft Agent Framework; not recommended for new projects. See this site's AutoGen page |
| huggingface/smolagents | 10k+ | A minimalist framework whose core is only about a thousand lines, built around CodeAgents (the model writes code as its action); for learners who want to see what an Agent really is |
| google/adk-python | 10k+ | Google's Agent Development Kit; the first choice in the Gemini ecosystem, with A2A protocol support |
| pydantic/pydantic-ai | 10k+ | From the Pydantic team, type safety first; for engineers who dislike framework magic and want a FastAPI-like experience |
TypeScript / JavaScript
- mastra-ai/mastra — the most complete Agent framework in the TS ecosystem right now, with workflows, memory, and evals built in; for frontend/full-stack teams.
- langchain-ai/langgraphjs — the JS version of LangGraph, concept-aligned with the Python release.
- google/adk-js — the TypeScript implementation of Google ADK.
Platforms and Low-Code
- Flowise — a drag-and-drop Agent builder (50k+ stars), for people who do not write code or need to validate an idea quickly; complex scenarios will hit a ceiling.
- ByteDance's Coze (the open source version of Coze) and Dify are the two most widely used open source Agent/LLM application platforms in China. This site has case studies for both Coze and Dify.
Examples and Templates
- Shubhamsaboo/awesome-llm-apps — 100+ runnable Agent/RAG application templates (120k+ stars), from travel assistants to multi-Agent teams, every one cloneable and runnable. Start here for portfolio inspiration, then rework what you find using this site's portfolio project guide—do not submit them as-is.
Ecosystem Components
- modelcontextprotocol/servers — the official collection of MCP reference servers (filesystem, Git, databases, and more). Study their structure before writing your own server.
- mem0ai/mem0 — an open source memory layer for Agents (50k+ stars). Before building long-term memory, read its paper and implementation to decide whether to build your own or adopt it. Pairs with this site's memory chapter.
Advice on reading source code
To truly understand the Agent Loop, close-reading one minimal implementation beats reading the docs of ten frameworks. The smolagents core is roughly a thousand lines—an afternoon's read—and once you finish, you will know what a "framework" actually does for you and what it hides. Big frameworks like LangGraph will make far more sense afterward.
5. Benchmarks and Datasets
The point of reading benchmarks is not to memorize scores—it is to understand how this field defines "hard." Below are the most frequently cited benchmarks in Agent evaluation; see this site's evaluation framework for details. First, a quick-reference table:
| Benchmark | What It Tests | In One Line |
|---|---|---|
| SWE-bench Verified | Resolving real GitHub issues | The de facto standard for Coding Agents |
| τ-bench / τ²-bench | Multi-turn conversation + tool calling | The reliability litmus test for customer-service / conversational Agents |
| GAIA | Web browsing + multimodal + multi-step reasoning | The comprehensive exam for the "general assistant" |
| WebArena / OSWorld | Web / desktop GUI operation | The proving ground for browser and computer-use Agents |
| Terminal-Bench | Long-horizon terminal tasks | A complement covering system-level capabilities |
| AgentBench | Combined evaluation across eight environments | An academic, multi-dimensional capability profile |
| BrowseComp | Locating information on the open web | Browsing challenges that are easy to verify yet hard to solve |
Details on each benchmark:
- SWE-bench — evaluates Coding Agents on real GitHub issues, and its Verified subset is the current de facto standard. Between 2024 and 2025, the resolve rate of top systems climbed from roughly 40% to over 80% (per Anthropic's engineering blog)—the single best indicator for tracking Coding Agent progress.
- τ-bench / τ²-bench — a conversational Agent benchmark from Sierra that simulates multi-turn tool-use scenarios such as retail and airline customer service. It specifically tests pass^k (getting it right every time) rather than merely pass@1 (getting lucky once).
- GAIA — a general-assistant benchmark jointly launched by Hugging Face, Meta, and others (paper: arXiv:2311.12983). Its 466 human-designed questions test the combined abilities of web browsing, multimodal understanding, and multi-step reasoning; for a time it was the default reference point in "general Agent" marketing.
- WebArena — a self-hosted web task environment (e-commerce, forums, collaboration software) that measures browser Agents' end-to-end completion rates. Its sibling OSWorld extends the battleground to the entire desktop OS.
- Terminal-Bench — long-horizon tasks in a terminal environment (compiling kernels, training models, and the like), covering the system-level capabilities SWE-bench does not touch.
- AgentBench — a multi-environment benchmark from Tsinghua's THUDM group (eight environments including operating systems, databases, the web, and card games), useful for internalizing the idea that "Agent capability is not a single dimension."
- BrowseComp — OpenAI's browsing benchmark, with questions that are "easy to verify, hard to solve," specifically testing an Agent's ability to find needles on the open web.
Discipline for reading leaderboards
Scores for the same benchmark under different harnesses (scaffolds) are not directly comparable—Anthropic says so explicitly in its eval article: "evaluating an Agent really means evaluating the harness + model combination." Before trusting any score, ask three questions: what scaffold was used, how many runs and which metric (pass@1 or pass@k), and at what cost. Use leaderboards to observe trends, not to make purchasing decisions.
6. Communities and Information Sources
Newsletters and Podcasts
With information sources, quality beats quantity. These three cover both "technical depth" and "industry perspective," and all of them are free:
- Latent Space — the "AI Engineer" newsletter and podcast run by swyx and Alessio; one of the most technically dense English sources, and the usual first outlet for in-depth Agent interviews.
- Import AI — Jack Clark's (Anthropic co-founder) weekly newsletter, running for years; it leans toward policy and research trends and helps you build a big-picture view.
- Simon Willison's Weblog — an engineer's running log on LLMs, with hands-on tests and commentary on new models and tools; he rarely misses a major release in the Agent ecosystem.
Communities
- Reddit: r/LocalLLaMA (frontline model and inference discussion), r/AI_Agents (Agent-focused), and r/ClaudeAI (plenty of hands-on posts from the Claude ecosystem).
- Discord: the official Discord servers of open source projects such as LangChain, Hugging Face, and CrewAI are direct channels for asking questions and tracking roadmaps—better suited than issue trackers for "is this the right way to do it" questions.
- Hacker News — the first venue where major releases and classic posts get discussed; authors often show up in the comments. The 12-Factor Agents mentioned earlier first gained traction here.
- X/Twitter: follow the official accounts (@AnthropicAI, @OpenAI, @GoogleDeepMind) plus a small number of engineer-type voices (such as @_akhaliq's daily paper updates); the more accounts you follow, the worse the signal-to-noise ratio. One enforceable rule: if your following list exceeds 50, prune it.
Chinese-Language Communities
- The WeChat official accounts Synced, QbitAI, and AI Era (Xinzhiyuan) — the "big three" of Chinese AI news (Synced's website, QbitAI's website). Fast and timely, good for skimming headlines and tracking developments; filter the deeper technical judgments yourself—their role is media, not engineering reference.
- WaytoAGI — the most active open source AI knowledge-base community in the Chinese-speaking world, maintained as Feishu (Lark) docs, with a rich stock of Chinese-language Agent tutorials and case studies.
- Juejin's AI section and Zhihu's "LLM" topic — the main producers of Chinese-language engineering write-ups. Quality varies widely, so filter by author rather than by platform.
7. Awesome Lists
Use these when you need exhaustiveness rather than curation (for instance, when researching whether a niche category already has a ready-made wheel). The right way to use an awesome list is to scan it by category and pick 2-3 candidates for a PoC—not to click through from top to bottom:
- e2b-dev/awesome-ai-agents — the most diligently maintained catalog of Agent projects (about 30k stars), split into open source and closed source, still actively merging PRs as of 2026.
- kyrolabs/awesome-agents — another high-star Agent directory, categorized by framework, platform, and application.
- dair-ai/Prompt-Engineering-Guide — the best-known awesome repo in prompt engineering (70k+ stars), with indexes of papers and guides. Use alongside this site's prompt engineering page.
- ai-boost/awesome-harness-engineering — a niche directory that emerged in 2026, collecting articles and tools on harness/scaffold engineering (agent loop, evals, memory, permissions, observability)—closer to the engineering front line than generic Agent lists.
- For papers, go straight to this site's paper map and core papers list—more Agent-focused than any general awesome list.
One last honest word
A resource list cannot solve the "bookmarked but never read" problem. A more effective strategy than growing your bookmarks bar: settle on 2-3 first-hand sources (say, Anthropic Engineering + Latent Space + one Chinese-language source), finish them every week, then reproduce the patterns you read about in your own small Agent. Keeping the input-to-output ratio at 1:1 is the fastest way to learn in this field.
8. How to Use This List by Role
The same list deserves a different approach depending on where you stand:
| Who You Are | Read First | Skip for Now |
|---|---|---|
| Career changer starting from zero, still building fundamentals | Microsoft / Hugging Face courses → Building effective agents → this site's concept clarifications | Benchmark details, multi-Agent papers |
| Working engineer shipping within a project | Official docs + Cookbooks → 12-Factor Agents → Manus context engineering → the cost and observability chapters | Course-style content (too slow) |
| Preparing for job interviews | All the classic blog posts + the design thinking behind SWE-bench/τ-bench + this site's interview questions and knowledge map | Low-code platforms |
| Making product/technical decisions | OpenAI's enterprise guide + the cost analysis in Anthropic's multi-Agent retrospective (multi-Agent burns about 15x the tokens) + framework selection | Source-code-level material |
An executable four-week rhythm, for reference:
- Week 1: Read Building effective agents plus the first half of any beginner course, and hand-write a minimal Agent with an official SDK—no frameworks.
- Week 2: Read the two context engineering pieces from Manus and LangChain, then go back to your own Agent and make it KV-cache friendly; start reading the Claude Cookbooks notebooks relevant to your direction.
- Week 3: Read Demystifying evals and write a small 20-case eval set for your own Agent (Anthropic's advice is precisely to start small), and read 12-Factor Agents as a checklist against your work.
- Week 4: Pick one framework (LangGraph or the OpenAI Agents SDK) and rewrite your Week 1 Agent to feel what the framework actually saves you from; meanwhile pick a template from awesome-llm-apps and hack it into a portfolio project.
After four weeks, the information-source resources in this list (Section 6) take over—your main job shifts from "systematic learning" to "continuous following."
References
- Building effective agents — Anthropic Engineering — the most-cited foundational article in Agent engineering (2024-12); the baseline for this page's "classics" selection.
- Context Engineering for AI Agents: Lessons from Building Manus — the Manus team's hands-on summary of context engineering (2025-07); original text verified via FetchURL.
- How we built our multi-agent research system — Anthropic Engineering — multi-Agent architecture and token economics analysis (2025-06); verified via FetchURL.
- Demystifying evals for AI agents — Anthropic Engineering — Agent evaluation methodology (2026-01); verified via FetchURL. The source of this page's grader taxonomy and pass@k/pass^k concepts in the benchmarks section.
- A practical guide to building agents — OpenAI — OpenAI's official guide to building Agents (2025).
- Model Context Protocol official documentation — the first-party MCP spec and the versioned specification.
- microsoft/ai-agents-for-beginners — Microsoft's free beginner Agent course (released in 2025, with an official Chinese translation).
- e2b-dev/awesome-ai-agents — the reference for this page's inclusion criteria in the awesome lists section; star-count orders of magnitude come from its repo page and third-party star statistics (mid-2026).