Skip to content

Career Guide: The Job Landscape

At a glance A panorama of the AI job market — role-by-role comparison of seven job types (LLM algorithm engineer, RAG/Agent application engineer, prompt engineer, AI product manager, AI infra engineer, data & eval engineer, AI safety engineer), a role-to-content mapping table, market trends, and how to use the four career pages.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Career Guide: The Job Landscape ​

Hold on to this one sentence: the career section solves two problems — first, seeing clearly which roles exist in the AI hot-concepts space and what each demands; second, closing your personal gap against them. This page is the navigation map for both: the role landscape gives you market information, the trend section helps you read direction, and the content map tells you which four pages to dig into.

The most expensive sequencing mistake in a job hunt is "effort first, positioning later": you grind three months of Transformer interview questions, then discover the target role never tested any of it. The AI field moves especially fast — there was no "prompt engineer" job title at the end of 2022; by 2025 it shows up in job descriptions everywhere. So this section pulls you back on track: do the information work first, then the learning work.

1. What This Section Is: Read the Market, Then Prepare Yourself ​

The five pages in this section (this overview + four content pages) form a job-hunt pipeline where each stage's output feeds the next:

Your situation: you want a job / role change / move into the AI hot-concepts space
         │
         ▼
┌─────────────────────────────────────────────┐
│ (1) READ THE MARKET ── [The JD List](/career/jd-list)
│     What roles are companies hiring for?     │
│     Which requirements show up in JD after   │
│     JD? Hard requirements vs nice-to-haves?  │
└─────────────────────────────────────────────┘
         │
         ▼
┌─────────────────────────────────────────────┐
│ (2) READ YOURSELF ── [Knowledge Breakdown](/career/knowledge-map)
│     Map each JD skill word to "can / can't / │
│     half-can", build a personal study list   │
│     instead of re-reading everything         │
└─────────────────────────────────────────────┘
         │
         ▼
┌─────────────────────────────────────────────┐
│ (3) PACKAGE YOURSELF ── [Resume Analysis](/career/resume-analysis)
│     Rewrite project experience through an    │
│     interviewer's eyes: model choices, eval  │
│     methods, failures and lessons learned    │
└─────────────────────────────────────────────┘
         │
         ▼
┌─────────────────────────────────────────────┐
│ (4) PROVE YOURSELF ── [Interview Question Bank](/career/interview-questions)
│     High-frequency questions + answer        │
│     frameworks — self-test at the end, not   │
│     on day one                               │
└─────────────────────────────────────────────┘

This sequence is the main line of the learning paths: start from the end and work backwards — derive what to learn from JDs and interview questions, aiming for full coverage of high-frequency test points rather than completeness. If you only have two or three weeks, follow this main line; if you have more time, interleave it with the site's regular route: What Are AI Hot Concepts? → core concepts → case studies → hands-on practice.

How this section relates to the rest of the site

The career section does not re-teach knowledge. It does two things only: market information and job-hunt method. For concepts, go to the concept landscape to sort out "AI vs ML vs deep learning vs generative AI", and to the glossary for definitions. The most common mistake at this stage is spending time on "one more chapter" instead of "organizing what you already know" — the career section exists for the latter.

2. The Role Landscape: Seven Roles at a Glance ​

Scan the summary table first, then take each role apart. The table below covers the seven most common AI hot-concepts roles in the Chinese and international job markets as of dataAsOf (2025-06), generalized from public job descriptions — it does not point to any specific company's open positions.

RoleCommon English TitleOne-Line PositioningCore Output
LLM Algorithm EngineerLLM Algorithm EngineerThe "model builder": pretrains, fine-tunes, aligns modelsMore accurate, more stable, better-behaved models
RAG/Agent Application EngineerRAG / Agent EngineerThe "integrator" wiring general models into product scenariosRAG Q&A and agent applications that run real business flows
Prompt EngineerPrompt EngineerThe "translator" designing prompts and interaction patternsPrompt templates that reliably reproduce high-quality output
AI Product ManagerAI Product ManagerThe person defining "what can AI actually solve"Shippable product requirements and evaluation plans
AI Infra Engineer (Inference)AI Infra / Inference EngineerMakes models fast and cheap to runLow-latency, low-cost inference services
Data & Eval EngineerData & Eval EngineerFeeds the model's data and holds the yardstickHigh-quality datasets and evaluation benchmarks
AI Safety EngineerAI Safety EngineerInsures the model and cages the risksSafety evaluations, red-teaming, governance plans

How to read the salary numbers (read this first)

The salary figures in this section are coarse-grained ranges compiled at writing time (dataAsOf: 2025-06) from public postings on recruiting platforms and third-party statistics such as levels.fyi. They are order-of-magnitude references for choosing a direction, not precise quotes. "Tier-1 Chinese cities" means full-time annual salary in Beijing, Shanghai, Shenzhen, Hangzhou, etc. (in RMB, covering common levels); "abroad" mostly means US tech companies (in USD). Pay for the same title varies several-fold by degree, experience, company, and interview performance; treat anything older than a quarter as a trend, not a quote. Actual offers and the latest statistics are authoritative.

1. LLM Algorithm Engineer — the model builder ​

Responsible for pretraining, continued training, instruction tuning (SFT), and alignment (RLHF/DPO) of large language models. Typically found on big tech foundation-model teams, research labs, and open-source model teams. This is the highest-barrier of the seven roles: pretraining positions usually require large-scale training experience (thousand-GPU class and up) or a relevant research background; new grads usually enter through fine-tuning, data, or evaluation tracks.

2. RAG/Agent Application Engineer — the model user ​

The fastest-growing role since 2023. The core job is wiring foundation models into concrete business scenarios: building Retrieval-Augmented Generation (RAG) knowledge-base Q&A, constructing agent workflows, handling tool calling and multi-step tasks. It doesn't require training models, but it does require you to know the model's boundaries precisely and combine it with engineering.

3. Prompt Engineer — the model translator ​

The standalone "prompt engineer" title has always been controversial, but it genuinely exists in the hiring market — and its more common form is a composite skill inside other roles. The job translates business needs into instructions models execute reliably: system prompt design, few-shot example construction, chain-of-thought (CoT) steering, structured-output constraints.

  • Core skills: prompt engineering methodology, instruction writing and debugging, output parsing and fallbacks, and an evaluation loop (any claimed improvement must be backed by data).
  • Typical pay: standalone roles appear mostly abroad and at AI-native companies; roughly RMB 250k–600k per year in tier-1 Chinese cities; roughly $120k–220k per year abroad.
  • Where to go on this site: Prompt Engineering is the main battlefield and The Prompt Engineering Playbook is the field manual; to prove you can actually do it, you also need to build LLM evaluations.

Will "prompt engineer" be a short-lived job?

A much-debated question. The pragmatic read: standalone prompt roles may shrink, but the era of "everyone writes prompts" amplifies the underlying skill — the typist role disappeared, yet typing became the default skill of every job. So learning prompting as a composite skill never loses.

4. AI Product Manager — the problem definer ​

The AI PM's core job isn't writing code; it's defining what AI can actually solve for users and how success is measured. Responsibilities include requirements analysis, model capability-boundary assessment, prompt and interaction-flow design, evaluation plans, and cost/compliance considerations. This is one of the most realistic entry points into AI for non-technical backgrounds — but it doesn't mean technical judgment is optional: PMs who don't understand model boundaries easily ship demos instead of products.

5. AI Infra Engineer (Inference) — the one who makes models affordable ​

"Works" and "works affordably" are two different things. AI infra engineers own inference serving, inference optimization and quantization, KV cache management, deployment framework choices, and GPU cost optimization. Demand for this role rose sharply after 2024, because the bottleneck of LLM applications shifted from "do we have a model" to "how much per hundred million tokens, and what's the P99 latency".

6. Data & Eval Engineer — feeding data, setting yardsticks ​

The bottleneck of LLM applications is shifting from "building models" to "building data and running evaluations", pushing data and evaluation roles from the periphery to the core: data cleaning and mixing, instruction data construction, human/model annotation systems, eval sets and benchmark building, red-teaming. It's the most underrated direction and the easiest place to accumulate differentiated experience — many LLM algorithm engineers start their careers doing exactly this.

  • Core skills: Python data processing, annotation guideline design, quality-control workflows, evaluation and benchmark methodology, RAG evaluation (retrieval quality + generation quality, two dimensions), data pipeline engineering.
  • Typical pay: roughly RMB 200k–500k per year in tier-1 Chinese cities; roughly $100k–200k per year abroad (senior eval engineers earn more).
  • Where to go on this site: LLM Evaluation and Benchmarks is the theoretical core and Build an LLM Evaluation Suite is the hands-on main line; Datasets and Tools Reference has a ready-made list of data resources.

7. AI Safety Engineer — insuring the model ​

Stronger models, bigger risks: prompt injection, jailbreaks, hallucinations, bias, privacy leaks. AI safety engineers own safety evaluation, red-teaming, content-safety filtering, and compliance implementation. In China these roles often sit inside content-safety departments; abroad they have matured into dedicated Responsible AI / Safety teams. See AI Safety and Governance for the underlying concepts.

  • Core skills: safety evaluation methodology, both sides of prompt injection and jailbreak attacks, content-safety policy, model bias assessment, compliance knowledge (data law, generative AI regulations, etc.).
  • Typical pay: roughly RMB 300k–700k per year in tier-1 Chinese cities; roughly $150k–280k per year abroad.
  • Where to go on this site: AI Safety and Governance is the home page; safety evaluation lands inside the framework of Evaluation and Benchmarks; to understand the attack surface you can't skip Alignment: RLHF and DPO.

Don't mistake a job title for an identity

The same title can mean completely different things at different companies: one company's "LLM algorithm engineer" writes SQL all day; another's "AI product manager" tunes prompts all day. Reading the JD always matters more than reading the title — which is why the second page of this section is The JD List, not an encyclopedia of titles.

3. Role Profiles: Modeling · Application · Engineering · Product ​

Put the seven roles into a four-dimension coordinate system and their personalities jump out. The four dimensions are:

            MODELING (pretraining / fine-tuning / alignment / quality)
                      ▲
                     /|\
                      │
      PRODUCT ────────┼──────── ENGINEERING
  (requirements /     │      (systems / deployment / data
   evaluation /       │       pipelines / infra / cost)
   communication)     │
                     \|/
            APPLICATION (prompts / RAG / agents / shipping)

Weights use a 1–5 scale (1 = rarely needed, 5 = core of the job):

RoleModelingApplicationEngineeringProductRole Personality
LLM Algorithm Engineer5231Scientist
RAG/Agent Application Engineer1542Application + engineering
Prompt Engineer1523Application + product
AI Product Manager1315Product
AI Infra Engineer (Inference)1251Engineer
Data & Eval Engineer2342Data + engineering
AI Safety Engineer2333Safety + governance

How to use this table

Don't just look for the biggest number — look at which column you can tolerate going deep on, long-term: modeling people can stand the tedium of tuning runs; engineering people enjoy the satisfaction of systems running stably; product people need the satisfaction of solving real user problems. Put "which kind of work do I want to do for years" ahead of "which role is hotter".

4. Role-to-Content Mapping Table ​

This table maps the knowledge each role demands to specific pages on this site. It's the most practical table on this page — once you've picked a target role, work down the "must-read pages" column and check them off; that's your personal learning path.

RoleMust-Read Pages (in priority order)Why These Pages
LLM Algorithm EngineerLarge Language Models (LLMs) → Transformers and Attention → Fine-Tuning and PEFT → Alignment: RLHF and DPO → LLM Evaluation and BenchmarksThe knowledge backbone of training and alignment roles; evaluation is the daily grind after onboarding
RAG/Agent Application EngineerRetrieval-Augmented Generation (RAG) → AI Agents → Vector Databases and Semantic Search → Build a RAG App from Scratch → Build an Agent from ScratchThe core mandate of application roles is "get RAG and agents right"
Prompt EngineerPrompt Engineering → The Prompt Engineering Playbook → LLM Evaluation and Benchmarks → Build an LLM Evaluation SuitePrompt value must be proven by evaluation — you need both
AI Product ManagerWhat Are AI Hot Concepts? → Prompt Engineering → LLM Evaluation and Benchmarks → Retrieval-Augmented Generation (RAG)Build the big picture first, then learn how to measure quality, then understand technical boundaries
AI Infra Engineer (Inference)Inference Optimization and Quantization → Deploying and Optimizing LLM Inference → Transformers and Attention → Large Language Models (LLMs)KV cache, quantization, batching all rest on understanding the architecture
Data & Eval EngineerLLM Evaluation and Benchmarks → Build an LLM Evaluation Suite → Retrieval-Augmented Generation (RAG) → Datasets and Tools ReferenceThe yardstick and the data sources for data roles live here
AI Safety EngineerAI Safety and Governance → Alignment: RLHF and DPO → LLM Evaluation and Benchmarks → Large Language Models (LLMs)Safety is the goalkeeper when alignment fails — you must understand alignment

The cross-role common denominator

Whatever the role, the Python ecosystem and the glossary are default skills; Pitfalls and Anti-Patterns is everyone's pre-flight checklist. Fill the common denominator first, then follow the table into your specialization.

The trends below are based on public JD observations and industry reports as of dataAsOf (2025-06). They are directional judgments, not precise statistics — re-verify against the latest information before citing.

1. "Knowing how to use LLMs" is moving from bonus to baseline ​

Unlike previous years, in 2025 LLM keywords no longer appear only in "LLM jobs" — ordinary algorithm, data, and even product JDs now list "familiar with mainstream LLM applications". Two consequences:

  • Universal-skilling: calling APIs, wiring up a simple RAG, and writing prompts are becoming office-and-engineering skills everyone is expected to have, sinking in like Office did;
  • Differentiation moving up: when everyone can use LLMs, resume differentiation shifts from "can you use it" to "how deep, how stable, how quantified" — evaluation ability, cost awareness, and end-to-end shipping become the new dividing lines.

2. Role specialization: three diverging tracks — training vs application vs infrastructure ​

The umbrella term "LLM jobs" is splitting into three very different career tracks:

TrackRepresentative RolesBarrierSupplyOutlook
TrainingPretraining/alignment researchersVery high (large-scale training / top conferences)ScarceConcentrated in a few foundation-model teams; fiercely competitive but a deep moat
ApplicationRAG/Agent engineers, AI product managersMedium (usage + engineering + evaluation)PlentifulThe largest volume of roles, and still growing as the tech platforms
InfrastructureInference optimization, MLOps, data/eval engineeringMedium-high (systems engineering)ModerateFastest-rising demand — everyone can call a model; few can run it stably and cheaply

3. Three directional judgments for job seekers ​

  1. Don't lock yourself into a single job title. Titles drift ("NLP engineer" → "LLM application engineer" → "agent engineer"); the skill stack is the constant. Get each of the four skill lines — RAG, agents, evaluation, inference optimization — to "can build it, can explain it", and you're hard to beat.
  2. Evaluation skill is the most underrated moat. As generation quality converges, evaluation, data, and scenario adaptation become the differentiation — which is exactly the logic behind rising data/eval roles. See LLM Evaluation and Benchmarks.
  3. Safety and governance are the long-term sure bet. The more widespread models become, the more structural the demand for AI safety talent; supply currently lags far behind demand, and the technical bar is relatively friendly.

6. How to Work Through This Section ​

Once you've picked a target role, go through this section in four steps. Deliverables and time references:

StepPageDeliverableSuggested Time
(1) Read JDsThe JD List2–3 target roles shortlisted; hard requirements and nice-to-haves separated1 day
(2) Map skillsJD Knowledge BreakdownPersonal "can / can't / half-can" list + study priorities1–2 days
(3) Benchmark resumeWhat Your Resume Should HighlightProject experience rewritten against the JD2–3 days
(4) Drill questionsInterview Question BankHigh-frequency checkpoint self-test + answer frameworksThe week before interviews

The order is the main line, not a suggestion

The four pages run "JD → skills → resume → drilling", each page's output feeding the next. Skipping around costs you most of the value — above all, don't charge into the Interview Question Bank on day one; without the positioning from the first two steps, drilling is just performing effort for yourself.

Further Reading ​

Keep reading on this site

References ​