Appearance
JD Knowledge-Point Breakdown
The previous page, the JD Landscape, showed you who's being hired; this page solves the next problem: for every line like "familiar with XXX" or "has XXX experience" in a JD, what exactly does the interviewer want to dig out of you?
JD wording is a compromise between HR and the hiring manager—phrasing is highly formulaic, but the intent behind it varies widely. The same "familiar with RAG" may mean one shop wants someone who can call an embedding API and another wants someone who can explain rerank and the evaluation chain. This page breaks down the high-frequency keywords one by one, each with the trio of test intent, pass bar, and self-verifying artifact—the artifact is the point, because "familiar with" costs nothing on a resume; only artifacts can tell the real from the fake.
Data basis
The keyword frequencies and requirement levels on this page synthesize Chinese-community aggregations of 50+ real AI application development JDs from the first half of 2026, plus English-community LangChain/Agent interview question banks from 2026; searched 2026-08. Time-sensitive data like salaries and version numbers should be re-verified with the latest searches.
1. High-Frequency Keywords, One by One
1. "Familiar with LangGraph / LangChain"
Test intent: frameworks aren't the point; the interviewer is confirming two things—can you use the framework correctly, and can you route around it when its abstractions get in the way. JDs that only say "can use LangChain" are mostly junior application roles; those that say "familiar with LangGraph" usually belong to teams with real orchestration complexity.
Pass bar:
- Junior: use LangChain 1.0's
create_agentto build a tool-calling agent within half an hour, and explain its execution loop (model → tool call → results written back → call the model again). - Mid-level: draw LangGraph's StateGraph model—state schema, nodes, conditional edges, reducers; know that
create_react_agentmigrated with 1.0 andlanggraph.prebuiltwas deprecated (the official LangChain/LangGraph 1.0 release of 2025-10-22 did this restructuring and promised no breaking changes before 2.0). - Senior: explain the checkpointer's persistence semantics—a state snapshot per super-step (
MemorySaverfor development,PostgresSaverfor production), and how that enables "resume from the breakpoint after a server restart" plusinterruptfor human-in-the-loop; name a scenario where you'd skip the framework and hand-write the loop, with the reason.
Typical follow-up: "What's the relationship between LangChain and LangGraph?"—the standard answer: since 1.0, LangChain's create_agent runs on the LangGraph runtime; LangChain provides the high-level abstractions and middleware (HITL approval, summarization, PII redaction), and LangGraph provides the underlying durable orchestration. Answering "LangGraph is the part of LangChain that draws graphs" takes you out of the running.
Self-verifying artifact: a GitHub project using create_agent + custom middleware; or a retrospective post comparing LangGraph vs a hand-written Agent Loop. See this site's LangGraph page and Agent Loop.
One judgment
The 2026 reality: most JDs say "familiar with LangChain/LangGraph," but what interviewers actually care about is whether you can explain the agent's execution model without the framework. Framework APIs churn every six months (0.x's AgentExecutor exited at 1.0); the underlying loop doesn't. People who memorized APIs get washed out by version churn; people who understand the loop don't.
2. "Has RAG project experience"
This is the most inflated line in any JD—and the easiest place to open up a real gap. Interviews explicitly split it into shallow and deep tiers:
Shallow tier (entry bar):
- Can walk the full pipeline: document loading → chunking (semantic/fixed-length, overlap strategy) → embedding → vector store (FAISS/Milvus/PGVector, any one) → top-k retrieval → assembling into the prompt for generation.
- Has actually built a runnable pipeline, even just a few hundred lines.
- Knows the embedding model and the generation model are two different things, and can name the selection criteria for mainstream embedding models (language, dimension, cost).
Deep tier (bonus material, and the default expectation for mid-level and above):
- Can discuss retrieval-quality tuning: hybrid search (BM25 + vectors), rerank (cross-encoders or dedicated rerank models), query rewrite / HyDE, and how chunking strategy affects recall.
- Can discuss evaluation: recall@k / MRR on the retrieval side; faithfulness and answer relevance on the generation side; knows frameworks like Ragas exist and their limits.
- Has scars and can recount them: semantic breaks in long-document chunks, lost tables and images, index-rebuild strategy after knowledge-base updates, multi-tenant knowledge isolation.
- Knows the applicability boundaries of RAG vs long context and RAG vs fine-tuning (see this site's RAG page).
One-line distinction: the shallow tier explains "how I hooked it up"; the deep tier explains "how I know it retrieves accurately and answers correctly—and how I fixed it when it didn't." If you can't answer "how did you handle bad cases in your RAG system," your experience credit drops to zero on the spot.
Self-verifying artifact: a RAG project with evaluation data—even an eval set of just 50 hand-labeled questions. As long as you show the baseline → optimization → metric improvement arc, it's more credible than 90% of resumes. Evaluation methods in Evals in Practice.
Below is a minimal recall@k self-check script that depends on no eval framework. When an interviewer asks "is your retrieval actually accurate" on the spot, this is what you should pull out:
python
# eval_retrieval.py — measure retrieval recall with 50 hand-labeled questions
import json
# each line of eval_set.jsonl: {"question": "...", "gold_doc_ids": ["doc_12", "doc_37"]}
# gold_doc_ids are the documents you labeled by hand as "must be retrieved to answer this question"
eval_set = [json.loads(line) for line in open("eval_set.jsonl", encoding="utf-8")]
def recall_at_k(k: int) -> float:
hits, total = 0, 0
for item in eval_set:
retrieved = retrieve(item["question"], top_k=k) # swap in your own retrieval function
retrieved_ids = {doc.id for doc in retrieved}
gold = set(item["gold_doc_ids"])
hits += len(gold & retrieved_ids) # required documents that were retrieved
total += len(gold)
return hits / total if total else 0.0
for k in (1, 3, 5, 10):
print(f"recall@{k} = {recall_at_k(k):.3f}")The point isn't the script (twenty minutes to write)—it's putting a before/after number like "after adding rerank, recall@5 went from 0.61 to 0.78" into the README. That's the line between deep and shallow.
3. "Understands Agent architecture"
Test intent: this is the prelude to a system-design question, testing whether you can draw a sensible component breakdown for a fuzzy "build an agent that can do X."
Pass bar—here's what the standard answer looks like:
┌─────────────────────────────────────────────┐
│ Agent Loop │
│ observe → think(LLM) → act(tool) → observe │
└──────────┬──────────────────┬────────────────┘
│ │
┌─────▼─────┐ ┌──────▼──────┐
│ Context │ │ Tool Layer │
│ (prompt + │ │ (function │
│ memory) │ │ call / MCP) │
└─────┬─────┘ └──────────────┘
┌─────▼─────────────────────────┐
│ Governance layer: eval / trace / guardrail │
└───────────────────────────────┘Your answer must cover four things: the single-loop model (observe-think-act, see What Is an AI Agent); the planning trade-off (ReAct thinking-while-doing vs Plan-and-Execute planning-then-doing, and which tasks suit which, see Planning); memory layering (in-session short-term vs cross-session long-term, see Memory Systems); failure handling (retrying failed tool calls, cutting off infinite loops, when humans step in, see Human-in-the-loop).
Bonus material: proactively discussing the applicability boundaries of multi-agent—"Multi-Agent is over-engineering in most scenarios; start with a single agent plus good tools; only go multi-agent when the task naturally divides and context isolation pays off." Saying this in a 2026 interview is a visible plus, because it reflects real engineering taste (details in Multi-Agent Architectures).
Self-verifying artifact: an independently written Agent design document (your course project's README will do), with the trade-offs of the four areas above spelled out.
4. "Prompt Engineering"
Test intent: by 2026 virtually no role titles anyone "Prompt Engineer"—it has become a baseline capability for all Agent engineers, like "knows SQL." When this line appears in a JD, the interviewer is confirming you can do prompts systematically, not tune them by feel.
Pass bar:
- Fundamentals: the layered structure of a system prompt (role/task/constraints/output format), the logic for choosing few-shot examples, and when structured output (JSON schema / tool-calling constrained decoding) is more reliable than "please output JSON."
- Engineering practice: prompt version management (prompts are code; they go in git with a change history), A/B and regression evaluation (prompt changes must be validated against an eval set, not eyeballed).
- Boundary awareness: knowing what prompts can't fix—missing knowledge needs RAG, unstable behavior needs few-shot or fine-tuning, insufficient capability needs a model swap. Says "let's try tuning the prompt first" but never treats the prompt as a panacea.
- Defense awareness: knowing prompt injection exists and its basic mitigations (see Security & Attack/Defense).
Self-verifying artifact: a prompt iteration record with evals attached—e.g. "v1 accuracy 62% → after adding negative examples of wrong outputs, v3 reached 89%, and the failure-sample distribution shifted from X to Y." Worth a hundred times more than the seven characters "prompt engineering expert." Material in Prompt Engineering and the Prompt Archive.
5. "Familiar with MCP"
This is a new requirement that only appeared en masse after 2025. MCP (Model Context Protocol) is an open protocol Anthropic introduced in late 2024, defining a standard interface between AI applications and external tools/data sources; by 2026 it has become the de facto standard for tool integration—"can write an MCP Server" is the new "can write a REST API."
Test intent: not protocol memorization—confirming you understand "why tool integration needed standardizing," and whether you can actually write a working server.
Pass bar:
- Concept layer: explain MCP's host/client/server three-party architecture; explain that it solves the M×N problem—M models/apps and N tools no longer need pairwise adapters.
- Hands-on layer: write an MCP Server with the official SDK (Python or TypeScript) exposing at least one tool and one resource, and actually invoke it through MCP Inspector or Claude/Claude Code.
- Judgment layer: know the boundaries between MCP and function calling, OpenAPI, and Agent Skills—MCP is the execution layer (wiring tools in), Skills are the orchestration layer (teaching domain flows to the agent); the two compose rather than replace each other—a high-frequency discrimination question in 2026 interviews.
- Security awareness: know the supply-chain risks of integrating third-party MCP Servers (a malicious server can induce the agent to perform dangerous operations), see Security & Attack/Defense.
Self-verifying artifact: an open-source MCP Server people actually use (even one solving your own small pain point); or a substantive PR to a well-known MCP Server repo. Details in Tools & MCP.
A minimal working Python MCP Server looks like this (the official SDK's FastMCP style, and a common starting point for whiteboard questions):
python
# server.py — exposes one tool, callable by any MCP client (Claude Code, your own agent)
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("issue-tracker") # server name; clients identify it by name
@mcp.tool()
def search_issues(keyword: str, limit: int = 5) -> list[dict]:
"""Search issues by keyword. The docstring becomes the tool description, directly shaping the model's call decisions."""
# in a real project this queries a database / calls an internal API; never swallow exceptions in a tool—
# return readable error text on failure so the model gets a chance to self-correct
return db.query(keyword)[:limit]
if __name__ == "__main__":
mcp.run() # defaults to stdio transport, good for local agents; use streamable-http for remote servicesTwo details you only learn by writing one, which map exactly to the hands-on test points: a tool's docstring is part of the prompt—write it vaguely and the model will pass garbage arguments; and error handling should translate exceptions into natural-language returns rather than raising, giving the agent loop a chance to recover. Mention these two unprompted in an interview, and "familiar with MCP" is confirmed on the spot.
An honest note
Between "knows about MCP" and "has used MCP" lies one real hands-on build. This requirement's frequency is still climbing fast, and it's among the highest-ROI single skills of 2026—the protocol itself isn't complicated; one weekend is enough to produce a decent server.
6. Quick-Reference Table for Other High-Frequency Requirements
| JD phrasing | What it actually tests | Where on this site |
|---|---|---|
| Familiar with ReAct / Reflexion patterns | Paper-level concepts + knowing when they don't apply | Anatomy of an Agent |
| Multi-Agent development experience | Division of labor/communication/aggregation, and more so the judgment of "why not to" | Multi-Agent |
| Familiar with Dify / Coze / n8n | Building on low-code platforms + knowing their ceiling | Coze, Dify |
| Agent evaluation/observability experience | Trace analysis, eval set construction, production metrics | Evaluation, Observability |
| Solid backend/engineering skills | Async, concurrency, rate limiting, cost control—not AI knowledge but used daily | Cost & Performance |
| LLM fine-tuning experience | SFT data construction, when to fine-tune vs RAG/prompt | Concept Clarification |
Note the last row: many application-role JDs say "fine-tuning experience preferred," yet in practice a team may fine-tune a model once a year. That's JD-writing inertia—don't be scared off by a lack of fine-tuning experience, but be ready with a clear-eyed "here's when I would choose to fine-tune" answer in interviews.
7. A Prior Judgment: Algorithm Role or Application Role?
Before parsing keywords, confirm which type of JD you're reading; the two are graded on entirely different axes:
- Agent algorithm roles (titles like "Agent Algorithm Engineer," "LLM Algorithm Engineer"): mostly master's/PhD candidates; the test centers on NLP/deep-learning fundamentals, top-conference papers or Kaggle/ACM records, model training and evaluation methods. Most of this page's content is merely the entry ticket for those.
- Agent application/engineering roles (titles like "AI Agent Development Engineer," "AI Application Development Engineer"): the test center is exactly all the keywords in section 1 of this page—frameworks, RAG, architecture, prompts, MCP—plus solid backend engineering. This is the main battlefield for career changers and new graduates.
The test is simple: if a JD weighs "papers," "top conferences," and "model training" more heavily than "LangChain," "RAG," and "MCP," it's an algorithm role; otherwise an application role. 2026 Chinese-community JD aggregations show application roles outnumber and outgrow algorithm roles, and are more forgiving about academic background—which is exactly this page's assumed reader.
2. Skills Radar (Text Edition): Three Capability Profiles
Six dimensions profile three tiers of candidates: framework usage, RAG, architecture design, prompts, MCP/tools, and engineering governance (evals/observability/cost).
Scale: 0=unfamiliar 1=knows concepts 2=has built hands-on 3=production experience 4=can lead a team and set the approach
Junior (new grad/career change) Mid-level (1-3 yrs) Senior (3 yrs+ / architect)
Framework 2 can build and run 3 scarred, knows the detours 3 can select—and can say "don't use"
RAG 1-2 pipeline running 3 tuning + eval experience 4 led a production retrieval system
Architecture 1 can recite the standard 2-3 designed an agent solo 4 trade-offs of complex systems
answer
Prompt 2 structured fundamentals 3 eval-driven iteration 3 prompts are just one tool
MCP/tools 1 knows the protocol 2-3 wrote their own server 3 architecture calls on the tool ecosystem
Governance 0-1 heard of evals 2 has run offline evals 3-4 production observability + cost loopThe key difference between tiers isn't "knows more," but the level of judgment:
- Junior: knows what each component is and can snap the parts together. The typical weakness: every problem gets the same solution (reaching for a ReAct loop on any requirement), no evaluation awareness, and "it runs" is the finish line.
- Mid-level: can independently ship an agent feature with real users. The signature is a library of failure cases—when talking about a project, they volunteer "we originally did it this way, hit problem X, and changed to Y." Evaluation and observability have moved from "heard of" to "done."
- Senior: the core skill is subtraction. Can judge "this requirement doesn't need an agent—rules suffice," "this doesn't need multi-agent—a single agent plus good tools is enough," "this feature isn't worth shipping—the unit economics don't clear." Also accountable for a system's long-term evolution: how to design state management to survive the next six months of requirement changes, and how to build the eval system to support continuous iteration.
Self-test method: take a project you actually built and ask, dimension by dimension, "can I tell a detail only a hands-on builder would know." Any dimension you can't answer scores 0-1, no matter how many articles you've read.
3. Learning Priority Order: By ROI
Assuming you already have Python/backend fundamentals and 2-3 hours a day, ordered by "interview pass-rate improvement per unit of time":
- Hand-write an Agent Loop (1 week). No framework; use the OpenAI/Anthropic SDK to implement the observe-think-act loop plus two or three tools. This is the foundation of everything that follows, and the best reverse lever against "familiar with framework" questions—every framework abstraction corresponds to a pain point you hit while hand-writing. See Build It Yourself.
- RAG project + evaluation loop (2-3 weeks). Covers the highest-frequency JD line. The key is pushing "it runs" to "it has metrics": build a 50-item eval set, run a baseline, do one or two optimization rounds, record the number changes.
- Get hands-on with LangChain/LangGraph 1.0 (1-2 weeks). With the first two steps done, the framework is just a different syntax. Focus on
create_agent, middleware, and LangGraph's checkpoint/interrupt; don't go back and learn the 0.x-era APIs (mixing old and new in an interview costs you points). Learning path in the Learning Path Overview. - Build an MCP Server (from one weekend). The protocol itself is simple; write one that solves a real need of yours and publish it to the community. This has extremely high marginal returns in 2026—not many people can do it, yet JDs asking for it keep climbing.
- Evals and observability (ongoing, interleaved). From your second project on, attach a trace tool (LangSmith / Langfuse / OpenTelemetry, any one) and build the habit of "run evals before changing anything." This is the dividing line between amateur and professional, and the scarcest interview narrative.
- Multi-Agent, planning strategies, and other advanced topics (as needed, 2-4 weeks). They're late not because they don't matter, but because judging when not to use them requires sufficient single-agent experience first. Learning multi-agent directly turns into a hammer looking for nails.
- Deep paper reading (long-term, 2-3 per week). Core papers like ReAct, Reflexion, and Toolformer are the ammunition depot for architecture questions, see Paper Reading Paths. It ranks last because it's slow to pay off but compounds—suiting long-term investment.
One principle throughout
Every step must leave a self-verifying artifact (a GitHub repo, an article with data, a PR), synced onto your resume. Knowledge acquired without an artifact counts as not acquired in the job market. How to package artifacts in Resume Analysis; project ideas in Portfolio Projects.
4. Using This Map Before an Interview
The 48-hour-before usage: copy the target company's JD line by line, compare against section 1's breakdown, and tag each line "can go deep / can discuss / must admit I can't." For the first two, prepare a 90-second project story each (context → your decision → quantified result); for the third, script an honest answer in advance—"I haven't gone deep there yet, but I understand its role is X, and I plan to fill that via Y" is far safer than bluffing.
The most dangerous moment in an interview isn't a knowledge gap—it's the panic after being caught. This map's value is that before you open your mouth, you already know your true level on every keyword. Common wording traps and responses in the Interview Questions.
Finally, a checkable self-audit list covering everything on this page:
| Check item | Pass marker |
|---|---|
| Can hand-draw the Agent Loop and explain each step's inputs/outputs | Whiteboards it within 5 minutes and can walk through a bad case live |
| Framework | Explains the LangChain/LangGraph 1.0 relationship and one deprecation change |
| RAG | Has an "improved metric from X to Y" optimization story |
| Architecture | For any requirement, produces a component breakdown + one explicit trade-off within 10 minutes |
| Prompt | Can show a prompt iteration record with version numbers and eval results |
| MCP | Owns a server they wrote that has been genuinely called |
| Governance | Can state the specific trace-tool usage in their project and one problem it surfaced |
References
- LangChain and LangGraph Agent Frameworks Reach v1.0 Milestones — the official 1.0 release notes, covering create_agent, middleware, persistence semantics, and the deprecation list; the authoritative reference for "familiar with LangGraph" questions.
- No algorithm grind, no papers: 2026 AI application development is hiring like crazy (with 50+ real JD analyses) — a first-hand aggregation of Chinese JD keyword frequencies and role distribution.
- The 2026 hiring market shifts: AI goes from bonus to baseline, requirements exposed by role — the actual usage contexts of keywords like LangChain/LangGraph, ReAct, and RAG in 2026 JDs.
- Top LangChain Interview Questions and Answers for 2026 (DataCamp) — a representative English-community question bank for framework interviews in 2026.
- What is the Model Context Protocol (MCP)? — the MCP official documentation entry point; first-hand material on the protocol architecture and SDKs.
- AI Agent interview questions part 4: MCP, Chrome DevTools, CDP session reuse — how Chinese community interviews actually probe MCP.
- ai-agent-engineer-handbook: JD Requirements — an open-source job-hunting handbook's observations on MVP project thresholds and the algorithm/application role split.