Appearance
Portfolio Projects
There's a harsh reality about resume screening for Agent roles: people who write "familiar with LangChain, understand RAG" are a dime a dozen, while people who can produce "here's a working system, here's the demo video, here's the eval report" are vanishingly rare. This page lays out 6 portfolio projects from shallow to deep, each designed to the standard of "doable in 1-2 weeks, quantifiable, and able to withstand follow-up questions." You don't need to do them all—completing 2-3 and polishing the presentation materials is enough to support a job search.
Before picking a project, be clear about what you're trying to prove. When employers evaluate Agent-track candidates, they're really looking at four things:
- End-to-end engineering ability: not calling an API once, but handling failures, retries, timeouts, and cost control.
- Evaluation awareness: is there an eval set, a pass-rate number, a bad-case analysis—the watershed between toys and products.
- Architectural judgment: why use or not use a given framework, why single-Agent instead of multi-Agent.
- Communication and presentation: is the README clear, the demo video, the quantified results.
The 6 projects map nicely onto these abilities.
1. Project Overview
| # | Project | Duration | One-line pitch | What it proves |
|---|---|---|---|---|
| 1 | Personal knowledge base Q&A | 1 week | Complete RAG pipeline + retrieval evaluation | Fundamentals |
| 2 | CLI coding assistant | 1-2 weeks | Hand-written minimal Agent Loop + tool system | Understanding of principles |
| 3 | Deep research assistant | 2 weeks | Multi-step search, report generation, traceable citations | Long-task orchestration |
| 4 | Browser automation Agent | 2 weeks | Booking/price-comparison scenarios, real web manipulation | Tools and robustness |
| 5 | Multi-Agent content factory | 2 weeks | Plan-write-review pipeline | Multi-Agent judgment |
| 6 | Vertical business Agent | 2 weeks+ | Support tickets/data analysis + human approval | Production thinking |
Difficulty rises roughly in order, but not linearly: project 2 is harder than project 1 in "understanding" rather than "engineering effort," and project 6 is hard in "restraint"—knowing where you must stop and hand control to a human.
Go deep on one, rather than shallow on six
Interviewers are unmoved by "I've built 6 projects" and intensely interested in "what was the hardest bug in this project and how did you track it down." Suggested combination: project 2 (proves you understand principles) + project 3 or 4 (proves you can build complex systems) + project 6 (proves production awareness). The rest are backups.
2. Project 1: Personal Knowledge Base Q&A (a complete RAG pipeline)
One-line pitch: turn your own notes/documents/saved web pages into a Q&A system with citation of sources, one that can say "I don't know" when it can't answer accurately.
This is the "Hello World" of the Agent field and the most frequently discussed topic in interviews. The value of building it isn't the system itself—it's stepping through every link of RAG with your own hands.
Core challenges
- The real impact of chunking strategy on quality: fixed-size, semantic, and structure-based splitting produce very different results—let the data speak.
- Retrieval quality evaluation: build a test set of 30-50 "question → expected matching document chunk" pairs and compute recall@k. This is the step most people skip, and precisely the one that earns the most credit.
- Hallucination control: when retrieval finds nothing relevant, does the model admit it or make something up? You need to handle this in the prompt and in fallback logic.
Tech stack: Python; embeddings from the OpenAI text-embedding-3 series or a domestic model (e.g. the BGE series); any vector store—Chroma/Qdrant/pgvector—at the portfolio stage the choice doesn't matter; any mainstream model for generation.
Milestones (1 week)
- Days 1-2: document parsing and splitting (one format each for Markdown/PDF is enough), vectorization and ingestion.
- Days 3-4: retrieval + prompt assembly + generation; command-line Q&A with cited chunks attached to answers.
- Day 5: build a 30-item eval set, run recall@k plus a human spot check of "answer correctness," record the numbers.
- Days 6-7: fix bad cases (tune chunk size, add hybrid search or rerank), write the README.
Presentable artifacts: in the README, a "before/after recall@k" comparison table, 3 Q&A screenshots, and 1 "refuses to answer" screenshot (proof of hallucination control). Demo video under 2 minutes.
Sample resume line:
Built a personal knowledge-base Q&A system (RAG) end to end: Markdown/PDF ingestion and semantic retrieval, a self-built 50-item eval set, and—by adjusting chunking strategy and adding rerank—lifted retrieval recall@5 from 62% to 89%, with traceable citations attached to every answer.
Likely interview follow-ups: how did you pick the chunk size? Why that vector store? How do you evaluate generation quality, not just retrieval quality? What do you do when a user's question contains multiple intents (hint: query decomposition)?
3. Project 2: CLI Coding Assistant (minimal Agent Loop + tools)
One-line pitch: with no Agent framework at all, hand-write a terminal coding assistant that can read files, edit code, run commands, and self-correct—200 lines of core code proving you really understand the Agent Loop rather than just calling frameworks.
This project packs far more punch on a resume than its line count suggests. The core of products like Claude Code and Cursor is a loop plus a set of tools (see the Claude Code teardown); write it once yourself and you can hold your own when the conversation turns to "ReAct" and "tool use."
Core challenges
- Tool protocol design: how should the tool schema be written so the model rarely mis-calls? When argument validation fails, how do you feed the error back to the model so it can self-correct?
- Termination conditions: the model can loop on tool calls forever. You need a triple safety net: max turns, token budget, repeated-action detection.
- Security boundaries: running shell commands and writing files must have an allowlist/confirmation mechanism, or your demo can blow up in front of the interviewer.
Tech stack: Python or TypeScript; call Anthropic/OpenAI's native tool-use API directly, no framework; a toolset of 4-6: read_file, write_file, run_shell, list_dir, search_code.
Minimal Agent Loop skeleton (Python + Anthropic SDK, current API shape as of 2026):
python
import anthropic
client = anthropic.Anthropic()
def run_agent(task: str, max_turns: int = 20):
messages = [{"role": "user", "content": task}]
for turn in range(max_turns):
resp = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=4096,
tools=TOOLS, # tool schema list
messages=messages,
)
messages.append({"role": "assistant", "content": resp.content})
# no tool calls = task complete
if resp.stop_reason != "tool_use":
return resp
# run the tools and feed the results back
results = [execute_tool(b) for b in resp.content if b.type == "tool_use"]
messages.append({"role": "user", "content": results})
raise RuntimeError("exceeded max turns, forced stop") # one of the three safety netsMilestones (1-2 weeks)
- Days 1-2: minimal loop running, supporting only
read_file+ conversation. - Days 3-5: round out the toolset, add dangerous-operation confirmation (print the command before
run_shelland wait for y/n). - Days 6-7: the self-correction loop—after a compile/test failure the model reads the error, edits the code, reruns; record a "fixing a bug end to end" video.
- Days 8-10 (optional): add an eval: find 10 real small tasks (e.g. "add unit tests for this function") and measure the first-try pass rate.
Presentable artifacts: a 3-minute terminal screencast (the model reads code, edits code, runs tests, and fixes errors by itself), a loop flow diagram in the README with every failure-handling point annotated. Emphasize "zero framework."
Sample resume line:
Implemented a CLI coding assistant from scratch (no Agent framework): hand-wrote the Agent Loop and the tool-use protocol, supporting 6 tools including file read/write and shell execution, with dangerous-operation confirmation and loop circuit-breaking; achieved a 70% first-try pass rate on 10 real coding tasks, with 80% of failures recovered via self-correction.
Likely interview follow-ups: how does it fall short of Claude Code (hint: context compaction, sub-Agents, permission system—link to your understanding of context engineering)? How did you design the error messages for when the model passes wrong tool arguments? How do you prevent it from deleting an important file?
4. Project 3: Deep Research Assistant (multi-step search + report generation)
One-line pitch: input a question, and it automatically decomposes sub-questions, searches in multiple rounds, cross-validates sources, and outputs a research report with per-claim citations—an open-source reproduction of the core Manus/Deep Research experience.
Since 2025, Deep Research products (OpenAI, Gemini, Perplexity, and Manus) have become the Agent field's most mainstream moment, and the open-source side has two mature references: GPT Researcher (~28k stars) and LangChain's open_deep_research (~12k stars). Building it yourself happens to cover four high-frequency interview topics: "planning, tool calling, long-form writing, citation tracing."
Core challenges
- Task decomposition and iteration: the first round of search results exposes new sub-questions, so planning can't be one-shot. Reference the plan-and-execute idea in Planning Patterns.
- Deduplication and conflict handling: when sources contradict each other, the report should present the disagreement, not force-pick a side.
- Traceable citations: every claim links to a specific URL + fetched snippet, letting readers click and verify. This is the only difference between "a research report" and "a long essay the model made up."
- Cost control: a single deep research run can burn hundreds of thousands of tokens, so cap search rounds/source counts (see Cost Optimization).
Tech stack: Tavily/Brave Search API for search (both have free tiers); LangGraph for orchestration (bonus framework experience) or pure Python; a long-context model for report generation; roll your own citation management (URL + fetch time + source snippet).
Milestones (2 weeks)
- Days 1-2: single-round version—question → search → summary with citations, proving the end-to-end path.
- Days 3-5: add a planner: decompose the question into a sub-question list, research each, then assemble a structured report (summary/body/reference list).
- Days 6-8: add a reflection loop: after each round of searching, judge "is the information enough"; if not, generate follow-up queries, at most N rounds.
- Days 9-11: citation tracing and conflict presentation; add a token-budget guardrail.
- Days 12-14: run 5 real questions end to end, human-check citation accuracy, write the README comparing answer quality of "asking the model directly" vs "the research pipeline."
Presentable artifacts: 2-3 full research reports as PDF/Markdown (pick questions with real informational value, e.g. "2026 vector database selection comparison"), an agent workflow architecture diagram, and citation-accuracy statistics (e.g. "spot-checked 60 citations, 58 matched the original").
Sample resume line:
Built a deep research Agent: plan-and-execute architecture with question decomposition, iterative search, and reflection loops, generating research reports with per-claim citations; an average task calls search 15 times and consumes ~80k tokens, with 96% citation accuracy (60 items human-checked).
Likely interview follow-ups: how do you decide "the research is done"? How do you filter SEO junk/AI-generated content from search results? Versus just using Perplexity, what's your differentiation? When sources contradict, how does your report handle it?
5. Project 4: Browser Automation Agent (booking/price-comparison scenarios)
One-line pitch: given one natural-language instruction ("compare the price of this book across three platforms"), the Agent opens the browser itself, navigates, fills forms, and extracts structured results—staring down the messiness of real web pages.
Browser operation is the hardest link in the Agent tool chain: pages are built for humans, not models, and pop-ups, CAPTCHAs, dynamic loading, and A/B redesigns can break the flow at any moment. Completing this project proves you can handle "real-world dirt." The mainstream path in 2026 is clear: use browser-use directly (the ~100k-star open-source library, the de facto standard in the Python ecosystem) or Stagehand (from Browserbase, TypeScript), or else attach browser capabilities to your own Agent through Playwright MCP.
Core challenges
- Perception representation: compress the DOM into a form the model can digest (accessibility tree, numbered interactive elements); raw HTML blows the context and performs poorly.
- Fragility and recovery: elements not found, page redesigns, network timeouts. Every step needs retries and a "take another path" fallback.
- Anti-scraping and compliance: target-site risk controls, login state, rate limits. At the portfolio stage, only touch public pages that permit automation, and don't show CAPTCHA bypassing in the README.
- Things that must not be done, don't do: real payments and real orders stop at the final step—a human confirms.
Tech stack: browser-use (Python) or Stagehand (TS), pick one—don't write DOM parsing from scratch; both are built on Playwright; the model needs strong vision/reasoning (screenshot understanding is more robust than raw DOM but pricier).
User instruction: "Check the price of item X on platforms A and B and give a recommendation"
│
▼
┌─────────────┐ retry on failure / new selector ┌──────────────┐
│ Agent Loop │ ◄───────────────────────────────── │ Browser │
│ (plan+reflect)│ ── click/type/scroll ───────────► │ (Playwright) │
└──────┬──────┘ └──────────────┘
│ structured record produced after each site
▼
┌─────────────┐
│ Aggregate + compare │ ──► JSON/Markdown report
└─────────────┘Milestones (2 weeks)
- Days 1-3: get browser-use's official examples running, swap in your own target scenario, complete a single-site flow.
- Days 4-6: multi-site task orchestration, unified structured output (JSON schema constraints).
- Days 7-9: hardening—retry strategy, timeouts, graceful degradation after page redesigns, screenshots archived on failure.
- Days 10-12: success-rate statistics (run the same task 20 times, record success rate and failure-cause distribution)—the most valuable number here.
- Days 13-14: record the demo, write the README, state the compliance boundaries explicitly.
Presentable artifacts: split-screen demo video (Agent reasoning log on the left, live browser on the right), success-rate table, a collection of failure-case screenshots ("how I handled these 5 failure types").
Sample resume line:
Built a cross-platform price-comparison Agent on browser-use: natural-language instructions drive the browser through search, filtering, and information extraction, outputting structured comparison reports; retry and degradation strategies lifted task success rate from 45% in the first version to 85% (20 repeated runs), averaging 90 seconds per task.
Likely interview follow-ups: how was the success rate measured, what sample size? What happens when the page changes (is there a self-healing mechanism)? Why browser-use instead of writing Playwright scripts yourself (the key to the answer: Agents are for uncertain tasks; fixed flows are more reliable as scripts—this judgment itself is a plus)? How do you handle login walls/CAPTCHAs?
6. Project 5: Multi-Agent Content Factory (plan-write-review pipeline)
One-line pitch: three Agents with distinct jobs—planner (topics and outlines), writer (drafts), reviewer (fact-checking and style)—form a content-production pipeline where humans only review the outline and the final draft.
The point of this project is not "multi-Agent is cool"—quite the opposite: after finishing it you should be able to articulate exactly when multi-Agent is over-engineering. A single Agent doing the three jobs in sequence also runs; multi-Agent's value lies in context isolation (the writer doesn't need to see the planner's search junk), independent iteration (review standards can be tuned separately), and parallelism (multiple topics written simultaneously). Being able to work through that accounting clearly impresses interviewers more than the pipeline itself. For a systematic discussion of orchestration patterns, see Multi-Agent Architectures.
Core challenges
- Handoff artifact design: what passes between Agents isn't chat logs but structured artifacts (outline JSON, drafts, revision lists); pin down the schema first.
- The reviewer Agent's credibility: "fact-checking" requires giving it search tools, otherwise it's just another model making things up again.
- Refuse infinite edit loops: the writer ↔ reviewer revision cycle must have a turn cap and a human as final arbiter, or the two Agents will edit each other until the end of time.
Tech stack: LangGraph for orchestration (still mainstream in 2026, with mature supervisor/pipeline patterns) or CrewAI (faster to pick up, see Framework Selection); a strong long-form model for writing; the reviewer Agent gets search tools.
Milestones (2 weeks)
- Days 1-3: single-Agent sequential version runs end to end (plan→write→review in one go), as the baseline.
- Days 4-7: split into three Agents, define the handoff schema, get the pipeline running.
- Days 8-10: add human approval points (outline confirmation, final-draft confirmation) and a revision turn cap.
- Days 11-14: comparative experiment—same batch of topics, single-Agent version vs multi-Agent version, human blind-quality rating + token cost comparison, conclusions written into the README. This comparison is your core insight.
Presentable artifacts: a pipeline architecture diagram, 2-3 complete outputs (with the full trail of outline, draft, review comments, final draft preserved), and a quality/cost comparison table of single-Agent vs multi-Agent.
Sample resume line:
Designed a multi-Agent content-production pipeline (plan/write/review division of labor, LangGraph orchestration): defined structured handoff protocols and human approval nodes; blind-rated quality improved 22% over the single-Agent baseline at 35% higher per-piece cost, with an analysis document on the applicability boundaries of multi-Agent based on the experiment.
Note that this line deliberately includes "35% higher cost"—a resume that dares to state trade-offs is far more believable than one that's all "improvements."
Likely interview follow-ups: why not one Agent with a three-part prompt? What if the reviewer objects but the writer disagrees (hint: arbitration mechanism)? Did you parallelize, and where's the bottleneck? If this went to production, which part would you replace first?
7. Project 6: Vertical Business Agent (support tickets/data analysis, with human approval)
One-line pitch: pick a real business scenario (support-ticket classification and draft replies, or natural-language data queries), get it to the level of "can demo to the business side," and put a human-approval gate in front of every high-risk action.
Of the 6 projects, this is the only one that tests "production thinking" more than "algorithmic tricks." The core tension of enterprise Agents is autonomy vs controllability: fully automatic and nobody dares use it, fully manual and it has no value. Your design—which actions run automatically, which need human sign-off, what the approval UI looks like, how approval records are kept—is the soul of this project. For a systematic treatment, see Human-in-the-Loop and Agent Security.
Core challenges
- Confidence tiers: high-confidence actions run automatically, low-confidence ones go to a human. How do you set the threshold? Calibrate with an eval set—don't guess.
- Approval experience: the approver shouldn't see the raw trace, but "what the Agent wants to do, on what basis, with one-click approve/edit/reject."
- Observability and accountability: every automatic execution and human intervention leaves a record (input, output, decision rationale, operator), see Observability.
Tech stack: using support tickets as the example—LangGraph for orchestration (its interrupt() + checkpointer mechanism is designed exactly for "pause, wait for approval, then resume"); 100-200 simulated ticket records you generate yourself; a minimal approval page for the frontend (Streamlit/Gradio is enough—don't burn budget on the frontend).
Milestones (2 weeks+)
- Days 1-3: data + baseline: ticket classification (intent recognition) + retrieval of similar historical tickets + draft replies.
- Days 4-6: eval-set calibration: measure classification accuracy and reply usability, set the confidence thresholds.
- Days 7-9: wire in the
interrupt()approval flow: low-risk actions (e.g. auto-tagging) execute directly; high-risk ones (e.g. sending a customer reply, refunds) pause for human approval. - Days 10-12: approval page + full audit logging.
- Days 13-14: write an "automation rate vs risk" analysis: what share of tickets can be fully automated, how many of the Agent's wrong decisions were intercepted.
Presentable artifacts: demo video of the approval flow (focus on the moment "the Agent makes a mistake and the approval gate catches it"—this says more about your understanding than 100 success cases), automation-rate/interception-rate numbers, and a one-page permission design document.
Sample resume line:
Built a support-ticket Agent (LangGraph + human approval flow): intent classification, similar-ticket retrieval, and reply drafting, with confidence thresholds calibrated on a 150-item eval set, achieving 68% fully automated handling; high-risk actions suspend via interrupt for human approval, intercepting 12 wrong Agent decisions during the trial run, with a complete auditable trail of every operation.
Likely interview follow-ups: how were the thresholds set, and how do you weigh false blocks against missed blocks? When an approver overrides the Agent, how does that feedback flow back to improve the system? (hint: this is the entry point to RLHF/online learning) What if the Agent and the user collude to trick approval? What would need to change to move this into finance/healthcare?
8. Turning Projects into Offers: Hard Requirements for Presentation
A finished project is only worth half its price; the other half is presentation. A few hard standards:
- The README is the storefront. Structure: one sentence saying what this is → a 30-second demo GIF/video → architecture diagram → how to run it → key design decisions and trade-offs → evaluation numbers → limitations. Limitations must be written; omitting them tells the interviewer you never thought about them.
- At least one quantified number per project. Success rate, accuracy, cost, latency—pick any, but it must exist, and the measurement method must be stated. "Works great" is filler; "85% success rate over 20 repeated runs" is evidence.
- Keep demo videos under 3 minutes, showing the complete loop (input → process → result), ideally including one failure and recovery.
- The code must run. Test the README's install steps on a clean machine; pin dependency versions; API keys via environment variables.
- Resume-to-interview continuity: each project gets 2-3 resume bullets, each = what you did + how + quantified result; the answers to interview follow-ups must be things you actually did in the project—never write numbers you can't defend. For overall resume strategy see Resume Analysis, and for systematic follow-up prep see Interview Questions.
The three most common ways to blow it
- Posting only a GitHub link, with no numbers and no demo—interviewers will not clone your repo.
- Demo videos showing only successes—any experienced interviewer assumes your success rate is 100% minus whatever you hid.
- Writing "proficient in LangGraph" on the resume but unable to explain
interrupt()'s recovery mechanism—the gap between "used" and "proficient" is exposed by one follow-up question.
One last suggestion: share a single "engineering chassis" across all projects—the same eval script template, trace recording format, README template. By the third project you'll find your efficiency doubled, and the chassis itself is engineering craftsmanship you can show an interviewer. If you'd rather get a minimal loop running before choosing projects, start with Build an Agent by Hand; to see how others turn a single project into a portfolio-grade retrospective, Common Pitfalls has plenty of cautionary tales.
References
- browser-use/browser-use (GitHub) — the mainstream open-source library for browser automation Agents; the first-choice foundation for project 4.
- langchain-ai/open_deep_research (GitHub) — LangChain's open-source deep research Agent; the reference implementation for project 3.
- assafelovic/gpt-researcher (GitHub) — another mature open-source deep research project, good for dissecting its multi-step search and citation design.
- Browser Agents in 2026: Browser Use vs Stagehand vs Skyvern vs Playwright MCP — a 2026 comparison of browser Agent toolchains, backing project 4's tech choices.
- LangGraph website — the orchestration framework for projects 5 and 6, with documentation on human-in-the-loop and persistent state design.
- AI Agent Frameworks Compared (2026) — a 2026 survey of the Agent framework landscape, used to confirm each framework's current status.
- LangGraph Human-in-the-Loop (2026): interrupt() tutorial — an API-level reference for project 6's approval flow.