Skip to content

Portfolio Projects: Three Resume-Worthy Builds

At a glance Three progressively harder agent portfolio projects for engineers moving into agent roles—a mini harness with a full evals suite, a production-grade MCP server, and a multi-agent / long-horizon task agent—each mapped to the most frequently requested job-description keywords, annotated with its technical points, acceptance criteria, and STAR resume bullets, with effort estimates at a light and a full tier.

Portfolio Projects: Three Resume-Worthy Builds ​

The skill benchmark page's conclusion: agent engineering roles want "people who can take an agent from demo to production—and prove it actually got better." Which raises the question: if you don't currently have that kind of experience on the job, what do you prove it with?

The answer is portfolio projects—but not just any "chatbot built with LangChain." These projects are reverse-engineered from job-description (JD) keyword frequency: each one maps onto several of the most frequently requested requirements, and each of those has to survive the interview grilling—"how did you do it, why that way, and what broke?"

Where the numbers come from

Every JD quoted on this page comes from a survey of 44 real, actively hiring postings conducted August 11–12, 2026 (27 roles at Chinese companies, 17 at international ones—a rough, hand-checked keyword count). The full list with sources is in JD Checklist. Every quote is attributed to a company and role; anything we couldn't verify against the original text was left out.

Choosing projects: let keyword frequency set the exam ​

Start by looking at what employers are actually paying for. Across the 27 China-based postings, the top of the frequency table reads: agent runtimes/harnesses (15 postings), tool calling (13), evals (12), context engineering (10); further down are multi-agent (9), observability (7), MCP (5), and memory and sandboxing (4 each). Among the 17 international postings, evals (10) tops the list too.

The three projects are designed straight off this frequency table:

text
           Three projects × JD keyword-frequency coverage
┌──────────────────┬───────────────────────────────────────────────┐
│ Project 1        │ Agent loop ● tools ● context ● evals          │
│ mini harness     │ (the top four keywords, all in one pass)      │
│  + evals         │                                               │
├──────────────────┼───────────────────────────────────────────────┤
│ Project 2        │ Tools/MCP ● auth & safety ● fundamentals      │
│ production-grade │ (MCP named in 5 JDs, sandbox/permissions      │
│ MCP server       │ named often; backend skills are the hidden    │
│                  │ hard gate in every JD)                        │
├──────────────────┼───────────────────────────────────────────────┤
│ Project 3        │ Multi-agent ● memory ● observability          │
│ multi-agent /    │ (9/4/7 postings; the differentiator for       │
│ long-horizon     │ architecture-heavy roles)                     │
│ agent            │                                               │
└──────────────────┴───────────────────────────────────────────────┘

One ground rule first

The difference between a portfolio piece and a toy is exactly one thing: has it been seriously evaluated, and has it handled real failures. Tencent Hunyuan's posting asks for "first-hand feel for the capability boundaries and failure modes of agentic coding"—feel can't be faked. Every project below lists "cause failures and document the failure modes" in its acceptance criteria; that's not optional, it's the soul of the project. A project that only runs the happy path, put on a resume, points the interview straight at your blind spots.

Project 1: A mini harness with a full evals suite ​

The real problem it solves. Agent demos are everywhere, yet nobody can answer "after this one-line system prompt change, did the agent actually get better or worse?" This project takes v1–v3 of the progressive tutorial (agent loop → tool system → planning and compaction) as its foundation, then adds a three-layer eval suite following Agent Evals from Scratch: every change to your own harness comes with numbers to back it up.

Which high-frequency requirements it covers. One project hits the top four of the frequency table:

  • Agent loop / harness (15 postings): Meituan's "Agent Harness Engineer" role asks for hands-on work on "core capabilities including system prompts, tools, skills, context management, the agent loop, task state, and exception recovery";
  • Tool calling (13 postings): Moonshot AI's "Agentic Growth Engineer" posting asks for the ability to "build high-quality scaffolding around an agent: how context is organized, how tools are abstracted, how loops terminate, how failures recover, how effectiveness is measured"—that sentence is practically this project's acceptance checklist;
  • Evals (12 China / 10 international): ByteDance's "Agent Harness Engineer (AI Data & Safety)" role asks for candidates who can "implement an automated eval pipeline, version regression, and A/B experimentation"; Cognition puts it even more bluntly: "Build evals that actually capture what matters... making sure the numbers mean something";
  • Context engineering (10 postings): truncation and session compaction are context engineering in its smallest, most concrete form.

Key technical points. A tool registry and a structured format for feeding errors back to the model; approval gates tiered by irreversibility; event-driven context compaction; three layers of assertions—unit, trajectory, outcome; pass rates reported over N runs per case; a mechanism for recycling bad cases into the eval set; CI wired up as a regression gate.

Two effort tiers.

TierScopeEstimated effort
LightGet v1–v3 running against a real model, plus 20 real cases and an eval script with three-layer assertions~1 week
FullLight tier + a CI regression gate + a multi-model / multi-prompt comparison matrix + a short experiment report on "what I changed and how the scores moved"~2 weeks

Acceptance criteria. The eval set has ≥20 cases and includes known successes and known failures; the grader has been calibrated against a MockLLM first; you can demo the full loop of "change one prompt → the eval scores move"; the README documents at least 3 failures you personally caused (tool errors stuck in a loop, context overflow, plan drift) with before/after trajectory comparisons.

How it reads on a resume (STAR).

"Built a three-layer eval suite for a self-built mini agent harness: a 20-case eval set, unit/trajectory/outcome assertions, and multi-sample pass-rate statistics (S/T); implemented tool-error classification and feedback, event-driven context compaction, and a CI regression gate (A); turned prompt iteration from gut feel into quantified regression on every change, surfacing 3 failure modes nobody had noticed before (R)."

Project 2: A production-grade MCP server ​

The real problem it solves. An agent's ceiling is often set on the tool supply side: the data the model wants isn't in its hands. MCP (Model Context Protocol) is today's de facto standard for tool integration, yet nearly every MCP demo you'll find online is hello-world grade—a wrapped public API, no error handling, no concept of permissions. An MCP server wired to a real data source, good enough to hand to other people fills exactly that gap.

Which high-frequency requirements it covers.

  • MCP (named by 5 China-based postings, including senior roles): Moonshot AI's "Senior Agent R&D Engineer" posting requires familiarity with the mainstream agent stack ("Skills, MCP, Sandbox, etc.") and real development experience on agent systems; Xiaohongshu's "Agent Harness Engineer" role asks for "deep interest or hands-on experience with LLM agents, tool calling, MCP, agent runtimes, coding agents, multi-agent systems and related directions"; Anthropic's Applied AI role lists MCP among its production LLM experience requirements;
  • The tools/security crossover: the same Xiaohongshu role asks for "building agent identity, authentication, and permission governance—solving privilege escalation, least-privilege, and security-boundary problems"; even Alibaba's internship posting already expects "engineering-grade approaches to risks like LLM hallucination and prompt injection"—an MCP server is the concrete artifact where all of these requirements become real;
  • The hidden item: engineering fundamentals. Retries, rate limiting, caching, logging, deployment—the backend skills sitting in the hard-requirements column of every JD.

Key technical points. Pick a data source you genuinely use (your own notes, an internal API, a public dataset); tool granularity and description writing (a tool description is a prompt aimed at the model); structured error returns instead of thrown exceptions; read/write operation tiering with authorization; timeouts, rate limiting, idempotency; wire it into Claude Code or any MCP client and test it for real.

Two effort tiers.

TierScopeEstimated effort
LightOne data source, 3–5 read-only tools, structured errors, end-to-end against one MCP client~1 week
FullLight tier + permission tiering and audit logs for write operations + rate limiting/caching/retries + an adversarial test set (injection-style parameters, unauthorized requests) + a deployment~2 weeks

Acceptance criteria. A real MCP client can drive it through an end-to-end task; every error path returns model-readable structured information (never a raw stack trace); write operations sit behind permission gates with audit records; the README has a threat-model section: what this server exposes to the model, and where the boundary is drawn.

How it reads on a resume (STAR).

"Designed and built a production-grade MCP server for XX data source, letting agents query and operate XX directly (S/T); designed tool granularity and descriptions, structured error feedback, read/write permission tiering with audit logs, and ran adversarial tests with injection-style parameters and unauthorized requests (A); the server is used daily by XX agent workflows, and tool-call error rate fell from XX% in the first version to XX% (R)."

Fill in real numbers of your own—if you have no real users, write "powers my own N daily agent workflows." Honest beats pretty.

Project 3: A multi-agent collaboration demo or a long-horizon task agent ​

The real problem it solves. A single agent on a long task hits two walls sooner or later: the context fills up, and the goal drifts. Pick one of two directions: a multi-agent collaboration demo (an orchestrator dispatches tasks to specialized subagents and merges their results, handling conflicts), or a long-horizon task agent (a job that takes dozens to hundreds of steps—"watch this source and produce a daily summary report"—held together by persistent memory and resumable runs).

Which high-frequency requirements it covers.

  • Multi-agent (9 postings): Moonshot AI's "Agent Product Engineer (multi-agent)" posting is the most instructive—"design the interaction protocols for agent-to-agent collaboration, treating collaboration itself as a new kind of harness"—and its hard requirements state that you can "quickly hand-roll a harness to validate agent collaboration and orchestration efficiency"; Baidu's "Agent Algorithm Engineer" role asks for "short- and long-term memory management and multi-agent coordination"; a Tencent Yuanbao role lists "multi-agent collaboration patterns, human-in-the-loop";
  • Observability (7 postings): ByteDance's harness role asks for "an end-to-end agent observability system: execution trace tracking, structured log collection, performance monitoring, anomaly alerting, visual debugging, trace replay"; Xiaohongshu asks for "full coverage across Trace/Log/Metric/Event, with debug replay and failure diagnosis"—a multi-agent system without traces is a black box, so this is a must-have, not a nice-to-have;
  • Memory (4 postings): the core of the long-horizon direction, mapping onto the persistence and externalization design in Memory Systems.

Key technical points. Task contracts for subagents (what goes in, what comes out, what happens on timeout); context isolation between orchestrator and workers—subagents bring back conclusions, not process (the core motivation for subagents); cross-session state persistence and resumable runs; end-to-end traces where every call from every agent can be replayed and attributed.

Two effort tiers.

TierScopeEstimated effort
Light1 orchestrator + 2 specialized subagents, structured task contracts, replayable JSONL trace logs~1.5 weeks
FullLight tier + persistent memory and resumable runs + trace visualization (even a bare-bones local web page) + a documented set of collaboration failure modes (lost tasks, conflicting results, delegation loops)~2–2.5 weeks

Acceptance criteria. It completes a task a single agent handles poorly (use Project 1's eval approach to compare single-agent vs. multi-agent completion rates—the most persuasive data you can bring); any run can be replayed in full and answer "which agent went off the rails at which step"; at least 2 multi-agent-specific failure modes and their countermeasures are documented.

How it reads on a resume (STAR).

"On XX long-horizon task, a single agent's completion rate sat at just XX% due to context bloat (S); I designed an orchestrator + specialized subagents architecture: structured task contracts, context isolation, end-to-end traces, and resumable runs (T/A); completion rate on the same task rose to XX%, and any failure can be replayed and attributed to a specific agent's specific step within 5 minutes (R)."

Choosing among the three projects ​

Project 1: mini harness + evalsProject 2: MCP serverProject 3: multi-agent / long-horizon
Keyword coverageThe top four (loop/tools/context/evals)MCP, permissions & safety, backend fundamentalsMulti-agent, memory, observability
Who it fitsEveryone—this one is mandatoryEngineers moving from backend into agent workAnyone with an algorithms/ML background or aiming at architecture roles
Suggested orderBuild firstBuild secondBuild last, reusing Project 1's evals

If you only have two weeks, do the full version of Project 1; with a month, add Project 2. Project 3 can wait—a seriously evaluated single agent beats a multi-agent system nobody knows works, which has been this site's design stance all along (see Harness Design Principles and Common Pitfalls & Anti-Patterns).

Further reading ​