Skip to content

Curated Resource List

At a glance A navigation guide to Agent learning resources organized into seven categories—official docs, books and courses, engineering blogs, open source frameworks, benchmarks and datasets, community sources, and awesome lists—each entry carrying a one-line verdict on its value and who it suits. All links verified.

This page contains time-sensitive content; data is current as of 2026-08. Job listings, pricing, and product features may have changed — verify against the original sources before citing.

Curated Resource List ​

The goal of this list is not to be exhaustive—it is to make sure every entry is worth your time. There are only three inclusion criteria: is the content first-hand (official docs, the author's own blog, the original repo), has it been tested in real engineering practice (production case studies, runnable code), and is it still maintained (substantive updates after 2025). An entry must meet at least two of the three to make the cut.

The list is organized as a learning path: start with official docs and the classic posts to build judgment, then pick a framework and get your hands dirty, then calibrate your understanding with benchmarks, and finally rely on information sources to stay current. If you are not sure what to learn first, work through the learning paths page and come back to follow this list step by step.

How to use this list

Information in the Agent field has a half-life of about 6 months. This list carries a dataAsOf stamp (2026-08); framework star counts and course availability may have changed, so check the original pages before quoting specific numbers. More important than bookmarking links: make a fixed habit of reading 2-3 first-hand sources every week—it is far more efficient than skimming second-hand news.

1. Official Documentation and Cookbooks ​

Official documentation is the only source in this field that will not mislead you—vendors are accountable for their own APIs. Third-party tutorials frequently copy from outdated API versions; official docs do not. The three vendors each have a distinct style, so pick what fits: Anthropic's docs lean toward design philosophy and best practices, OpenAI's toward guides and sample code, and Google's toward a structured reference manual.

Anthropic ​

  • Claude official documentation — the authoritative reference for the Agent SDK, tool calling, and MCP integration. Note that the domain has moved from docs.anthropic.com to docs.claude.com, so it is time to update old bookmarks. For anyone developing against Claude-family models.
  • Claude Cookbooks — an officially maintained collection of copy-and-run code snippets covering RAG, tool use, and Agent patterns, all runnable as-is. Ideal for hands-on learners who learn by adapting working code.
  • Claude Code documentation — even if you never plan to use Claude Code, its docs are the best observational sample of what configuration surface a mature Coding Agent exposes. It reads even better paired with this site's Claude Code case study.

OpenAI ​

  • OpenAI Agents SDK documentation — official docs for the lightweight Agent framework, with clean definitions of handoffs, guardrails, and tracing. See this site's OpenAI Agents SDK page for details.
  • OpenAI Cookbook — the official examples repo. The examples/agents_sdk/ directory contains complete notebooks spanning single-Agent setups to multi-Agent collaboration, including a side-by-side guide for migrating from the Claude Agent SDK—worth reading when you are weighing the two SDKs against each other.
  • A Practical Guide to Building Agents — OpenAI's enterprise guide to building Agents, published in 2025 (PDF). It breaks an Agent down into three elements—model, tools, and instructions—and offers criteria for when to evolve from a single Agent to multi-Agent designs. Written for technical decision-makers, but engineers get plenty out of it too.

A taste of the official SDK's minimalist style (adapted from the quickstart in the official docs; treat the docs as authoritative for details):

python
from agents import Agent, Runner

# Define an Agent: name, instructions, and available tools
agent = Agent(
    name="assistant",
    instructions="You are a rigorous assistant: state the conclusion first, then the evidence.",
)

# Runner drives the Agent loop; run_sync is the synchronous entry point
result = Runner.run_sync(agent, "Summarize this week's LLM news for me")
print(result.final_output)

Ten lines or fewer and you have a working Agent—which is why the post-2025 debate shifted from "should you use a framework" to "should you use a heavyweight framework." Lightweight SDKs have already compressed boilerplate to nearly zero.

Google ​

  • Agent Development Kit (ADK) documentation — Google's Agent framework, open sourced in 2025. The docs are top-tier by big-vendor standards, with a separate repository for each language implementation (Python/Java/Go/TS).
  • google/adk-python — the main ADK repo. Its samples directory deserves more attention than the docs themselves.

Protocol Layer ​

2. Books and Courses ​

Books ​

Few books devoted specifically to Agents hold up under scrutiny, so for now only one makes the cut as a foundation:

  • AI Engineering by Chip Huyen (O'Reilly, 2025) — strictly speaking this is not an Agent book but a panoramic engineering guide to building applications on foundation models. Even so, its chapters on evals, RAG, fine-tuning, and inference costs are precisely the foundations of Agent engineering. The companion repo chiphuyen/aie-book carries the author's curated resources. Best for engineers who want a systems perspective rather than a tour of frameworks. A Chinese translation is also in print—check the edition details on the bookstore page before buying.

    Suggested reading order: if you are short on time, read the evals and RAG chapters first—the twin cornerstones of Agent engineering. Skim the prompt engineering chapter, and leave fine-tuning until your business actually hits a bottleneck.

Why no more books

Agent engineering only began converging on a shared consensus at the end of 2024 (Anthropic's Building effective agents was published in December 2024), and the publishing cycle of a printed book cannot keep up with that pace. For now, blogs plus official docs plus papers are better vehicles; books are for shoring up foundations, not for chasing the frontier.

Free Courses ​

  • Hugging Face Agents Course — a free, open source Agent course with a certificate, built around smolagents to teach the Code Agent paradigm. The companion site lives under huggingface.co/learn, and a community-maintained Chinese translation is available. Ideal for a systematic start from zero.
  • Microsoft AI Agents for Beginners — Microsoft's beginner course, released in February 2025 and since expanded from 10 to 18 lessons. It covers design patterns such as tool use, Agentic RAG, planning, and multi-Agent systems, with an official Simplified Chinese version. Roughly 50k+ GitHub stars; a good fit for beginners with coding experience who can follow the lesson plan.
  • DeepLearning.AI: AI Agents in LangGraph — a short course taught personally by LangChain founder Harrison Chase (about an hour of material), moving from a hand-written ReAct loop to state and persistence in LangGraph. Free, and a good first contact with LangGraph.
  • The DeepLearning.AI short-course series as a whole deserves attention: the same platform hosts other Agent-related short courses such as Functions, Tools and Agents with LangChain. They are small, free, and handy for filling gaps in spare moments.
  • Google's 5-Day AI Agents Intensive, run with Kaggle — a live, bootcamp-style program centered on ADK, with course materials archived on Kaggle for self-paced study. A good fit if you want to stay on Google's stack.

About certificates

Certificates from these courses carry little weight in the job market—what interviewers want to see is what you built with the patterns the courses teach. Rework, extend, and deploy each course's capstone project and put it in your portfolio; that is worth far more than the certificate itself. See this site's resume analysis.

3. Classic Engineering Blog Posts ​

If papers tell you what is possible, engineering blogs tell you what actually works in production. The pieces below are the most-cited writings from 2024-2026 whose value has stood the test of time, ordered by how essential they are.

There is a method for reading this kind of article: do not just record the conclusions—record the boundaries within which they hold. For example, "multi-Agent beats single-Agent" only holds when tasks can be parallelized and the information exceeds a single context window; Anthropic says plainly in the same piece that most coding tasks are actually a poor fit for multi-Agent. Extracting these boundary conditions pays off in interviews and architecture reviews alike.

ArticleSource and DateValue in One Line
Building effective agentsAnthropic, 2024-12The most-cited engineering article in this field, bar none. Its core claim—"most scenarios do not need a complex framework; simple composable patterns suffice"—still holds today
Context Engineering for AI Agents: Lessons from Building ManusManus (Ji Yichao), 2025-07From KV-cache hit rates to "the file system as context," every conclusion here was paid for in production scars. Cited by this site's context engineering page and the Manus case study
How we built our multi-agent research systemAnthropic, 2025-06A retrospective on the multi-Agent architecture behind Claude Research, with hard conclusions such as "token consumption explains 80% of the performance variance." Read alongside multi-agent architecture
Effective context engineering for AI agentsAnthropic, 2025Turns context engineering from a slogan into a methodology: what belongs in the context and what should stay out
Context Engineering for AgentsLangChain, 2025Organizes context strategies into four types—write, select, compress, isolate—one of the most usable taxonomies around
Writing effective tools for agentsAnthropic, 2025A hands-on guide to tool design—writing tools for Agents and writing APIs for humans are two different crafts
Building agents with the Claude Agent SDKAnthropic, 2025-09The official guide to building with the same SDK that powers Claude Code. Its guiding idea—"hand it a harness with a working loop, not a pile of rules"—is highly representative. Pairs with this site's Claude Agent SDK page
Demystifying evals for AI agentsAnthropic, 2026-01A systematic guide to Agent evaluation: the three-way grader taxonomy, the difference between pass@k and pass^k, and how to evaluate each Agent type (coding / conversational / research / computer use). Read alongside this site's evaluation framework
12-Factor AgentsHumanLayer (Dex Horthy), 2025A production principles list with an anti-framework stance; entries like "own your context window" and "small, focused agents" have become industry mantras. Published as a GitHub repo—worth printing and pinning to the wall
Claude Code best practicesAnthropic, 2025-04Nominally about CLI usage, but in substance a mental model for collaborating effectively with a Coding Agent
Unrolling the Codex agent loopOpenAI, 2025The Codex CLI author personally dissects the Agent Loop implementation: how prompt prefixes align with the cache, and how to compact context once it exceeds the threshold. Read alongside this site's Agent Loop page
Effective harnesses for long-running agentsAnthropic, 2025-11How to make long-horizon Agents that run for hours execute reliably: harness design, checkpoint recovery, and cross-session memory. Long-horizon tasks are the main battleground of 2026

Two more blogs are worth subscribing to rather than reading piecemeal:

  • Anthropic Engineering — infrequent updates, but every post delivers. Half of the classics above came from here.
  • The LangChain blog — inherently promotional (it is their own framework, after all), but its "agent concept clarifications" and "production case interview" series are consistently good. Just keep facts and marketing separate as you read.

4. Open Source Frameworks and Repos ​

Star counts are order-of-magnitude figures from mid-2026, taken from public statistics pages. GitHub stars change constantly, so treat the repository homepages as authoritative. For the selection logic see this site's framework selection overview; this section is navigation only.

First, a take: the 2026 framework landscape has clearly stratified. The bottom layer is the vendors' official SDKs (lightweight, close to the API), the middle layer is orchestration frameworks (LangGraph, CrewAI—the kind that manage state and flow), and the top layer is low-code platforms. Most production teams actually choose either "official SDK plus a little homegrown orchestration" or "the full LangGraph stack," and frameworks straddling the middle are being squeezed. Before picking a framework, work out which layer your problem actually lives in.

Python Orchestration Frameworks ​

RepoStar CountPositioning and Fit
langchain-ai/langchain130k+The LLM application framework with the largest ecosystem and the most integrations; for when you need to quickly assemble ready-made components
langchain-ai/langgraphTens of thousands (roughly 30-50k)Graph-based, stateful Agent orchestration and a common default for production Agents; for teams that need fine-grained control over flow
openai/openai-agents-python10k+A lightweight SDK built around the handoff/guardrail/tracing trio; for getting started fast inside the OpenAI ecosystem
crewAIInc/crewai50k+Role-playing multi-Agent; prototypes in a few dozen lines of code. For business process automation and demos. See this site's CrewAI page
microsoft/autogen50k+Conversational multi-Agent framework. Note: the official repo is now in maintenance mode—Microsoft has shifted its focus to Microsoft Agent Framework; not recommended for new projects. See this site's AutoGen page
huggingface/smolagents10k+A minimalist framework whose core is only about a thousand lines, built around CodeAgents (the model writes code as its action); for learners who want to see what an Agent really is
google/adk-python10k+Google's Agent Development Kit; the first choice in the Gemini ecosystem, with A2A protocol support
pydantic/pydantic-ai10k+From the Pydantic team, type safety first; for engineers who dislike framework magic and want a FastAPI-like experience

TypeScript / JavaScript ​

  • mastra-ai/mastra — the most complete Agent framework in the TS ecosystem right now, with workflows, memory, and evals built in; for frontend/full-stack teams.
  • langchain-ai/langgraphjs — the JS version of LangGraph, concept-aligned with the Python release.
  • google/adk-js — the TypeScript implementation of Google ADK.

Platforms and Low-Code ​

  • Flowise — a drag-and-drop Agent builder (50k+ stars), for people who do not write code or need to validate an idea quickly; complex scenarios will hit a ceiling.
  • ByteDance's Coze (the open source version of Coze) and Dify are the two most widely used open source Agent/LLM application platforms in China. This site has case studies for both Coze and Dify.

Examples and Templates ​

  • Shubhamsaboo/awesome-llm-apps — 100+ runnable Agent/RAG application templates (120k+ stars), from travel assistants to multi-Agent teams, every one cloneable and runnable. Start here for portfolio inspiration, then rework what you find using this site's portfolio project guide—do not submit them as-is.

Ecosystem Components ​

  • modelcontextprotocol/servers — the official collection of MCP reference servers (filesystem, Git, databases, and more). Study their structure before writing your own server.
  • mem0ai/mem0 — an open source memory layer for Agents (50k+ stars). Before building long-term memory, read its paper and implementation to decide whether to build your own or adopt it. Pairs with this site's memory chapter.

Advice on reading source code

To truly understand the Agent Loop, close-reading one minimal implementation beats reading the docs of ten frameworks. The smolagents core is roughly a thousand lines—an afternoon's read—and once you finish, you will know what a "framework" actually does for you and what it hides. Big frameworks like LangGraph will make far more sense afterward.

5. Benchmarks and Datasets ​

The point of reading benchmarks is not to memorize scores—it is to understand how this field defines "hard." Below are the most frequently cited benchmarks in Agent evaluation; see this site's evaluation framework for details. First, a quick-reference table:

BenchmarkWhat It TestsIn One Line
SWE-bench VerifiedResolving real GitHub issuesThe de facto standard for Coding Agents
τ-bench / τ²-benchMulti-turn conversation + tool callingThe reliability litmus test for customer-service / conversational Agents
GAIAWeb browsing + multimodal + multi-step reasoningThe comprehensive exam for the "general assistant"
WebArena / OSWorldWeb / desktop GUI operationThe proving ground for browser and computer-use Agents
Terminal-BenchLong-horizon terminal tasksA complement covering system-level capabilities
AgentBenchCombined evaluation across eight environmentsAn academic, multi-dimensional capability profile
BrowseCompLocating information on the open webBrowsing challenges that are easy to verify yet hard to solve

Details on each benchmark:

  • SWE-bench — evaluates Coding Agents on real GitHub issues, and its Verified subset is the current de facto standard. Between 2024 and 2025, the resolve rate of top systems climbed from roughly 40% to over 80% (per Anthropic's engineering blog)—the single best indicator for tracking Coding Agent progress.
  • τ-bench / τ²-bench — a conversational Agent benchmark from Sierra that simulates multi-turn tool-use scenarios such as retail and airline customer service. It specifically tests pass^k (getting it right every time) rather than merely pass@1 (getting lucky once).
  • GAIA — a general-assistant benchmark jointly launched by Hugging Face, Meta, and others (paper: arXiv:2311.12983). Its 466 human-designed questions test the combined abilities of web browsing, multimodal understanding, and multi-step reasoning; for a time it was the default reference point in "general Agent" marketing.
  • WebArena — a self-hosted web task environment (e-commerce, forums, collaboration software) that measures browser Agents' end-to-end completion rates. Its sibling OSWorld extends the battleground to the entire desktop OS.
  • Terminal-Bench — long-horizon tasks in a terminal environment (compiling kernels, training models, and the like), covering the system-level capabilities SWE-bench does not touch.
  • AgentBench — a multi-environment benchmark from Tsinghua's THUDM group (eight environments including operating systems, databases, the web, and card games), useful for internalizing the idea that "Agent capability is not a single dimension."
  • BrowseComp — OpenAI's browsing benchmark, with questions that are "easy to verify, hard to solve," specifically testing an Agent's ability to find needles on the open web.

Discipline for reading leaderboards

Scores for the same benchmark under different harnesses (scaffolds) are not directly comparable—Anthropic says so explicitly in its eval article: "evaluating an Agent really means evaluating the harness + model combination." Before trusting any score, ask three questions: what scaffold was used, how many runs and which metric (pass@1 or pass@k), and at what cost. Use leaderboards to observe trends, not to make purchasing decisions.

6. Communities and Information Sources ​

Newsletters and Podcasts ​

With information sources, quality beats quantity. These three cover both "technical depth" and "industry perspective," and all of them are free:

  • Latent Space — the "AI Engineer" newsletter and podcast run by swyx and Alessio; one of the most technically dense English sources, and the usual first outlet for in-depth Agent interviews.
  • Import AI — Jack Clark's (Anthropic co-founder) weekly newsletter, running for years; it leans toward policy and research trends and helps you build a big-picture view.
  • Simon Willison's Weblog — an engineer's running log on LLMs, with hands-on tests and commentary on new models and tools; he rarely misses a major release in the Agent ecosystem.

Communities ​

  • Reddit: r/LocalLLaMA (frontline model and inference discussion), r/AI_Agents (Agent-focused), and r/ClaudeAI (plenty of hands-on posts from the Claude ecosystem).
  • Discord: the official Discord servers of open source projects such as LangChain, Hugging Face, and CrewAI are direct channels for asking questions and tracking roadmaps—better suited than issue trackers for "is this the right way to do it" questions.
  • Hacker News — the first venue where major releases and classic posts get discussed; authors often show up in the comments. The 12-Factor Agents mentioned earlier first gained traction here.
  • X/Twitter: follow the official accounts (@AnthropicAI, @OpenAI, @GoogleDeepMind) plus a small number of engineer-type voices (such as @_akhaliq's daily paper updates); the more accounts you follow, the worse the signal-to-noise ratio. One enforceable rule: if your following list exceeds 50, prune it.

Chinese-Language Communities ​

  • The WeChat official accounts Synced, QbitAI, and AI Era (Xinzhiyuan) — the "big three" of Chinese AI news (Synced's website, QbitAI's website). Fast and timely, good for skimming headlines and tracking developments; filter the deeper technical judgments yourself—their role is media, not engineering reference.
  • WaytoAGI — the most active open source AI knowledge-base community in the Chinese-speaking world, maintained as Feishu (Lark) docs, with a rich stock of Chinese-language Agent tutorials and case studies.
  • Juejin's AI section and Zhihu's "LLM" topic — the main producers of Chinese-language engineering write-ups. Quality varies widely, so filter by author rather than by platform.

7. Awesome Lists ​

Use these when you need exhaustiveness rather than curation (for instance, when researching whether a niche category already has a ready-made wheel). The right way to use an awesome list is to scan it by category and pick 2-3 candidates for a PoC—not to click through from top to bottom:

  • e2b-dev/awesome-ai-agents — the most diligently maintained catalog of Agent projects (about 30k stars), split into open source and closed source, still actively merging PRs as of 2026.
  • kyrolabs/awesome-agents — another high-star Agent directory, categorized by framework, platform, and application.
  • dair-ai/Prompt-Engineering-Guide — the best-known awesome repo in prompt engineering (70k+ stars), with indexes of papers and guides. Use alongside this site's prompt engineering page.
  • ai-boost/awesome-harness-engineering — a niche directory that emerged in 2026, collecting articles and tools on harness/scaffold engineering (agent loop, evals, memory, permissions, observability)—closer to the engineering front line than generic Agent lists.
  • For papers, go straight to this site's paper map and core papers list—more Agent-focused than any general awesome list.

One last honest word

A resource list cannot solve the "bookmarked but never read" problem. A more effective strategy than growing your bookmarks bar: settle on 2-3 first-hand sources (say, Anthropic Engineering + Latent Space + one Chinese-language source), finish them every week, then reproduce the patterns you read about in your own small Agent. Keeping the input-to-output ratio at 1:1 is the fastest way to learn in this field.

8. How to Use This List by Role ​

The same list deserves a different approach depending on where you stand:

Who You AreRead FirstSkip for Now
Career changer starting from zero, still building fundamentalsMicrosoft / Hugging Face courses → Building effective agents → this site's concept clarificationsBenchmark details, multi-Agent papers
Working engineer shipping within a projectOfficial docs + Cookbooks → 12-Factor Agents → Manus context engineering → the cost and observability chaptersCourse-style content (too slow)
Preparing for job interviewsAll the classic blog posts + the design thinking behind SWE-bench/τ-bench + this site's interview questions and knowledge mapLow-code platforms
Making product/technical decisionsOpenAI's enterprise guide + the cost analysis in Anthropic's multi-Agent retrospective (multi-Agent burns about 15x the tokens) + framework selectionSource-code-level material

An executable four-week rhythm, for reference:

  1. Week 1: Read Building effective agents plus the first half of any beginner course, and hand-write a minimal Agent with an official SDK—no frameworks.
  2. Week 2: Read the two context engineering pieces from Manus and LangChain, then go back to your own Agent and make it KV-cache friendly; start reading the Claude Cookbooks notebooks relevant to your direction.
  3. Week 3: Read Demystifying evals and write a small 20-case eval set for your own Agent (Anthropic's advice is precisely to start small), and read 12-Factor Agents as a checklist against your work.
  4. Week 4: Pick one framework (LangGraph or the OpenAI Agents SDK) and rewrite your Week 1 Agent to feel what the framework actually saves you from; meanwhile pick a template from awesome-llm-apps and hack it into a portfolio project.

After four weeks, the information-source resources in this list (Section 6) take over—your main job shifts from "systematic learning" to "continuous following."

References ​