Skip to content

JD List: LLM-Related Roles

At a glance JD templates and keyword radar organized by role type: responsibilities and common requirements across eight categories including Algorithm Engineer (LLM), NLP, Training, Inference Optimization, AI Applications, Agent, Evaluation, and Data Engineering. Includes a generic JD skeleton and a master list of high-frequency skill keywords.

This page contains time-sensitive content, current as of 2025-08; job descriptions, rankings, product features, and other information may have changed. Please verify with original sources before citing.

JD List: LLM-Related Roles ​

Remember this one-liner from this page: a JD is the cheapest available sample of "what the market wants" — reading a hundred JDs is more reliable than guessing ten times what an interviewer will test on. This page organizes JD templates by role type, extracting the common skeletons of responsibilities and requirements that repeatedly appeared in major companies' public JDs since 2023, and includes a "keyword radar" so you can see at a glance the skill weight for each role.

Important Disclaimer

All JDs on this page are template frameworks distilled from common patterns in public major-company JDs — they do not point to any specific company's open positions, and do not represent the actual JD of any particular company. Real JDs vary widely in responsibilities, requirements, years of experience, and education thresholds depending on the company, team, and level. Always check each company's official career page for the latest, most accurate information (data current as of dataAsOf: 2025-08).

1. How to Use This List ​

1. JDs Change — Skill Combinations Are the Norm ​

JDs in the LLM field evolve faster than in any other industry: in 2023, "proficient with ChatGPT API" was the hot phrase; in 2024 it became "familiar with RAG and vector databases"; by 2025, "Agent and evaluation experience" was added on top. Focus on the unchanging fundamentals in JDs (Transformer principles, Python, engineering skills), then focus on the shifting combinations (RAG, Agent, multimodal, quantization). Fundamentals determine whether you can survive across cycles; combinations determine what you're worth right now.

2. Distinguish "Hard Requirements" from "Nice-to-Haves" ​

Major companies' JDs follow a pattern in their wording:

WordingMeaningHow to Respond
Proficient / Deep understanding / Skilled atHard requirement; interviewers will likely dig deepPrepare a 30-minute deep-dive explanation
Familiar with / UnderstandSoft requirement; explaining the concept is enoughPrepare a conceptual answer + one example
Experience with XX preferred / BonusPreferred but not requiredAlign with it in your resume and provide evidence
Master's degree or above / Top conference papers preferredEducation and academic thresholdsBreak through with projects and portfolio if unmet

3. Three Actions When Reading JDs ​

  1. Circle skill keywords: Pull out every technical noun and cross-reference them with the Knowledge Breakdown;
  2. Level them: Mark each skill as "I know it / I'm shaky on it / I don't know it," then generate a study list;
  3. Store samples: Save 3–5 JDs for your target role each week. After a month, you'll see the trend and know which way the market is moving.

2. Generic JD Skeleton: Four Blocks Every LLM Role Shares ​

Regardless of the role, major companies' JDs almost always have this structure:

BlockCommon ContentDeeper Intent
ResponsibilitiesDesign/develop/optimize XX systems, keep up with cutting-edge tech, drive business adoptionInterview deep-dive area: What exactly did you do? What pitfalls did you hit?
Technical RequirementsPython/PyTorch, Transformer, fine-tuning, RAG, evaluation, deploymentInterview fundamentals area: maps to concepts and practice chapters on this site
Engineering RequirementsCode quality, distributed systems/serving, data processing, production stabilityTests "can you ship," not just "do you know concepts"
Soft SkillsSelf-driven, communication, cross-team collaboration, English literature readingUsually not strictly enforced, but project deep-dives will expose gaps

Treat a JD as a "mini exam syllabus"

A JD is essentially a mini exam outline: responsibilities = interview deep-dive questions, technical requirements = fundamentals/definitions, engineering requirements = system design questions. Reverse-engineering interview questions from JDs yields much higher hit rates than randomly reading interview experiences.

3. JD Templates by Role Type ​

Each category below gives: role definition → responsibilities → requirements (required / bonus). All are common-pattern templates.

1. Algorithm Engineer · LLM Direction ​

Definition: Select and compose LLM capabilities (prompts, RAG, fine-tuning) for business scenarios, and handle evaluation and iteration.

Responsibilities:

  • Design and deploy LLM solutions for business scenarios: scenario assessment, base model selection, prompt/fine-tuning/RAG strategy composition
  • Build business evaluation sets, establish offline and online evaluation systems, and quantify production impact
  • Keep up with the latest models and papers, assess their business applicability, and drive adoption
  • Collaborate with engineering teams to complete model serving and ensure production stability

Requirements:

  • Required: Proficient in Python; deep understanding of Transformer principles and fine-tuning methods (SFT/LoRA); familiarity with evaluation & benchmarks methodology; has delivered complete projects
  • Bonus: Has RAG or Agent production experience; familiar with alignment (RLHF/DPO); has online A/B testing experience

2. NLP Algorithm Engineer ​

Definition: Develop text understanding and generation capabilities, using both classical NLP and LLM methods.

Responsibilities:

  • Research and develop NLP capabilities: tokenization, entity recognition, text classification, semantic search, dialogue systems
  • Integrate LLMs (few-shot, fine-tuning, RAG) into existing NLP pipelines and assess impact
  • Build Chinese corpus processing and quality assessment pipelines
  • Participate in model evaluation and online performance monitoring

Requirements:

  • Required: Python and deep learning frameworks; familiarity with NLP fundamentals (tokenization, word embeddings, sequence labeling); understanding of language modeling and Transformers; has text-related project experience
  • Bonus: Familiar with Chinese pre-trained models and corpus characteristics; has LLM application (RAG/fine-tuning) experience; familiar with vector search

3. LLM Training Engineer ​

Definition: Own the training systems and data pipelines for pretraining/continued pretraining/post-training (SFT, RLHF/DPO).

Responsibilities:

  • Build training data pipelines: collection, cleaning, deduplication, mixing, quality monitoring
  • Configure and tune training frameworks and distributed parallelism (data/tensor/pipeline/expert parallel)
  • Monitor training dynamics (loss, gradients, throughput), diagnose and resolve non-convergence, overflow, and checkpoint recovery issues
  • Keep up with scaling laws and latest training methods, evaluate and experimentally validate them

Requirements:

  • Required: Solid deep learning fundamentals and training experience; familiar with PyTorch and distributed training (DeepSpeed/Megatron-LM); understand the full pretraining data pipeline; can independently troubleshoot training issues
  • Bonus: Has MoE training or large-model experience; familiar with CUDA and memory optimization; has CUDA/HPC background; published top-tier conference papers

4. Inference Optimization Engineer ​

Definition: Optimize LLM inference performance and serving, improving throughput while reducing latency and cost.

Responsibilities:

  • Research and deploy compression methods: model quantization (INT8/INT4), distillation, sparsification
  • Optimize inference frameworks (vLLM / TensorRT-LLM / SGLang, etc.) for operators, batching, and KV cache management
  • Conduct memory and performance analysis, stress testing, and tuning; build performance monitoring systems
  • Address business-side inference cost and latency optimization needs

Requirements:

  • Required: Deep understanding of the full Transformer inference pipeline (including KV cache and sampling); familiar with GPU architecture and CUDA programming; has performance analysis and tuning experience; proficient in Python and C++
  • Bonus: Familiar with mechanisms like PagedAttention and continuous batching; hands-on quantization experience (GPTQ/AWQ); familiar with the full deployment & serving chain

5. AI Application Engineer (LLM Applications) ​

Definition: Integrate general-purpose LLMs into business, handling the production of prompts, RAG, Agents, fine-tuning, and evaluation.

Responsibilities:

  • Develop LLM-powered business features: prompt design, RAG pipelines, tool calling
  • Build business evaluation sets and regression tests, quantifying the impact differences between models/prompt/retrieval strategies
  • Evaluate open-source models and frameworks, comparing the cost-effectiveness of self-deployment vs. API usage
  • Optimize end-to-end latency and cost; ensure production stability

Requirements:

  • Required: Proficient in Python; familiar with LLM APIs and open-source model deployment; understand the full RAG pipeline (indexing/retrieval/generation); has evaluation awareness and practice
  • Bonus: Has Agent development experience; familiar with vector databases and embeddings; has fine-tuning (LoRA) hands-on experience; familiar with engineering practices (CI, monitoring, observability)

6. Agent Engineer ​

Definition: Design and develop LLM-based agents: tool calling, planning, memory, multi-agent collaboration.

Responsibilities:

  • Design and implement Agent architectures: tool/function calling, task planning loops (ReAct, etc.), memory management
  • Build Agent reliability: error recovery, guardrails, cost control
  • Research and deploy Agent protocols and frameworks (function calling, MCP, etc.)
  • Design protection mechanisms against prompt injection and other security risks

Requirements:

  • Required: Python and systems engineering skills; understand LLM inference and prompt mechanisms; has experience with tool calling or workflow development; has complete Agent or automation product experience
  • Bonus: Has memory (long-term/short-term) and multi-agent design experience; has production stability experience with LLM-based Agents; familiar with vector search and evaluation

7. Evaluation Engineer ​

Definition: Build evaluation systems for model and product quality, providing the basis for model selection and iteration.

Responsibilities:

  • Build evaluation systems: public benchmarks + custom business sets + human/model judgment workflows
  • Run offline batch evaluations and regression tests to prevent degradation from model iterations
  • Design and analyze online evaluations (A/B testing, user feedback), quantifying production impact and risks
  • Address benchmark contamination and evaluation bias issues; maintain evaluation data quality

Requirements:

  • Required: Understanding of evaluation & benchmarks methodology; familiar with common benchmarks (MMLU/GSM8K/HumanEval, etc.) and evaluation tools (lm-eval-harness/OpenCompass); Python and data analysis skills
  • Bonus: Has LLM-as-a-judge production experience; has data labeling workflow management experience; familiar with hallucination and safety evaluation

8. Data Engineer · LLM Direction ​

Definition: Handle collection, cleaning, mixing, and quality assurance of pretraining/post-training corpora and business data.

Responsibilities:

  • Build large-scale corpus pipelines: collection, parsing, cleaning, deduplication, filtering, mixing
  • Produce instruction/preference data and ensure quality (including managing labeling teams and synthetic data)
  • Build data quality monitoring and compliance review mechanisms
  • Support the model team's data needs by providing reusable data infrastructure

Requirements:

  • Required: Python and data processing skills (Spark/Flink, etc.); familiar with deduplication (MinHash) and filtering methods; understand pretraining data principles; has data pipeline engineering experience
  • Bonus: Has synthetic data experience; understands copyright and compliance risks; has search/recommendation data processing experience

4. Keyword Radar: High-Frequency Skill Words by Role ​

The "frequency" ratings in the radar table are subjective assessments based on how often each skill appears across JD samples (●●● = high-frequency hard requirement / ●● = common / ● = bonus or team-dependent). Sample observations come from public major-company JDs between 2023–2025, and are for exam-preparation direction only.

RoleHigh-Frequency (●●●)Common (●●)Bonus (●)
Algorithm Engineer (LLM)Transformer, fine-tuning/LoRA, RAG, evaluation, Python/PyTorchPrompt engineering, Agent, alignmentMultimodal, papers, A/B testing
NLP Algorithm EngineerNLP fundamentals, Transformer, text classification/extraction, PythonVector search, fine-tuning, evaluationDialogue systems, knowledge graphs
LLM Training EngineerDistributed training, pretraining, data pipelines, PyTorch/DeepSpeedParallelism strategies, loss diagnostics, MoECUDA, papers, supercomputing experience
Inference Optimization EngineerQuantization, KV cache, vLLM/TensorRT-LLM, performance analysis, C++/CUDAContinuous batching, operator optimization, GPU architectureInference framework source contributions
AI Application EngineerLLM API, RAG, prompt engineering, Python, vector databasesAgent, LoRA fine-tuning, evaluationDeployment ops, CI/CD
Agent EngineerTool calling, ReAct, memory, Python, systems engineeringMCP, multi-agent, evaluationSecurity/injection defense, reinforcement learning
Evaluation EngineerBenchmark sets, evaluation methodology, lm-eval-harness, data analysisLLM-as-a-judge, regression testingManual labeling management, online A/B
Data Engineer (LLM)Data processing, cleaning & dedup, Spark/Flink, data pipelinesMinHash, filtering, complianceSynthetic data, labeling management

How to use the radar table

Treat the "●●●" column as mandatory interview territory, the "●●" column as "at least understand the principles," and the "●" column as "bonus on your resume if you have it, no loss if you don't." Cross-reference with the Knowledge Breakdown and check each item off.

5. Master List of High-Frequency Skill Words ​

Looking at all roles together, here are the highest-frequency skill words (coverage = number of role categories where it appears as a required skill):

Skill WordCoverageRelated Page on This SiteInterview Weight
Python8/8— (basic skill)Mandatory for coding rounds
Transformer/Attention7/8Transformer Architecture ExplainedCore mandatory
Fine-tuning (SFT/LoRA)6/8Fine-tuning / Fine-tuning PracticeHigh-frequency
Evaluation6/8Evaluation & Benchmarks / Evaluation PracticeHigh-frequency, easy to stand out
RAG5/8RAG PracticeHigh-frequency
Prompt Engineering5/8Prompt EngineeringHigh-frequency
Distributed Training3/8Framework & Tool SelectionMandatory for training roles
KV cache / Quantization / Deployment3/8Deployment & ServingMandatory for inference roles
Agent / Tool Calling3/8LLM-based AgentsFastest growing
Alignment (RLHF/DPO)3/8AlignmentDeep-dive topic

6. Reading Between the Lines: Filler vs. Substance in JDs ​

Not every sentence in a JD carries the same credibility. Here's how to distinguish common "filler" from "substance" in major-company JDs:

JD PhrasingReal MeaningPriority
"Good teamwork and communication skills"Generic boilerplate — everyone writes thisIgnore
"Strong self-drive and learning ability"Tech stack may not be defined yet; you'll be figuring things out on the jobReference only
"Understanding of mainstream LLMs and their applications"You've at least used them and understand basic principlesSubstance — prepare for principle-level questions
"Experience with LLM applications/fine-tuning/RAG projects"Need project evidence; interviewers will dig deepCore substance
"Familiar with distributed training frameworks"Hard threshold for training roles; mandatory fundamentalsCore substance
"Top conference publications / open-source contributions preferred"Differentiator; nice to have but not requiredBonus
"Can handle pressure / fast iteration pace"Signal for overtime and version rush deadlinesWarning flag

Three signals that a JD means "we're seriously hiring":

  1. Technical requirements are specific (naming specific frameworks, specific tasks) → the team knows what they want
  2. Clear deliverable descriptions ("responsible for XX system," "improve XX metric") → there's a real opening
  3. "Required" and "bonus" items are clearly tiered → they have mature talent standards

Conversely, JDs that are all vague boilerplate and loaded with "strong learning ability" are likely HR mass-postings — research the team's actual situation before applying, and don't waste time on interviews for positions that are probably just backup candidates.

7. Trend Observations (2023–2025) ​

  1. "Knowing how to use" is depreciating; "can prove effectiveness" is appreciating: Early JDs said "familiar with ChatGPT," now they say "capable of evaluation and regression testing." Interviews increasingly ask "how do you prove your solution works."
  2. RAG has moved from bonus to mandatory: Almost all application roles now require RAG production experience; "Agent" and "tool calling" requirements have been added since 2024.
  3. Training roles are consolidating at top companies: Pretraining-only roles exist only at a handful of model companies; most JDs have shifted to "fine-tuning/post-training on open-source models."
  4. Inference optimization has become its own role: Keywords like quantization, KV cache, and vLLM have risen dramatically in frequency over two years; candidates with systems backgrounds have grown more competitive.
  5. Data and evaluation roles have gone from "afterthoughts" to "first-class citizens": Increasingly many teams write data quality and evaluation systems into their JDs with formal leveling.

8. Quick Reference: English JD Keywords ​

For roles at foreign companies and global-facing positions, here are the most common keywords in English JDs (with Chinese equivalents and interview tips that map directly to testable topics on this site):

English KeywordMeaningChinese EquivalentInterview Tip
LLM / Foundation ModelLarge language model / foundation modelLarge language modelTransformer principles are mandatory
Prompt EngineeringPrompt designPrompt engineeringOften paired with in-context learning
Retrieval-Augmented GenerationRAGRetrieval-augmented generationFull pipeline is mandatory
Fine-tuning / PEFTFine-tuning / parameter-efficient fine-tuningFine-tuningLoRA principles are mandatory
LoRA / QLoRALow-rank adaptation and its quantized variantLoRARank, alpha, memory comparison
RLHFReinforcement learning from human feedbackAlignmentThree-step process is mandatory
DPODirect preference optimizationAlignmentCompare with RLHF
Agent / Tool Use / Function CallingAgent / tool callingAgentReAct loop and reliability
Evals / BenchmarkEvaluation / benchmarkEvaluationContamination and LLM-as-a-judge are bonus points
Context Window / Long ContextContext window / long contextLong contextRoPE extrapolation and RAG tradeoffs
QuantizationQuantizationInference optimizationINT8/INT4 tradeoffs
Inference OptimizationInference optimizationInference optimizationKV cache, batching, throughput
Latency / ThroughputLatency / throughputDeployment metricsTTFT/TPOT
Vector Database / EmbeddingVector DB / embeddingsRAG infrastructureSelection logic
Distributed TrainingDistributed trainingTraining engineeringDP/TP/PP/ZeRO
Hallucination / GroundingHallucination / groundingHallucinationCauses and mitigation
Production / MLOpsProduction environment / engineeringEngineering capabilityEvaluation and observability loops

Reading between the lines of English JDs

"Strong CS fundamentals" = you'll face tough coding challenges; "Experience shipping to production" = deep-dive into production stability and incident handling; "Familiar with research literature" = they'll test your paper reading and reproducibility skills. English JDs use subtler language — prepare for the testable topics behind each keyword.

9. Hard Thresholds: Education, Experience, and Equivalent Work ​

Thresholds are statistical patterns, not hard rules — but knowing the patterns helps you allocate your limited resume space wisely:

StageCommon Thresholds (common patterns across major-company public JDs)What to Do If Unmet
Fresh graduateMaster's is mainstream; algorithm/training roles generally require Master's or above; application/engineering roles accept Bachelor'sProve hands-on ability with runnable, complete projects + evaluation data
1–3 years"3+ years LLM/NLP/recommendation/search experience" appears frequentlyDeep projects can substitute for years: one thoroughly explained project ≈ one year of experience
3+ years / Senior"Independently responsible for complete systems," "experience with large-scale production"Focus interviews on production stability and post-incident retrospectives
PhD / Top conferenceHard threshold or strong bonus for training and research rolesFill paper gaps with open-source contributions and reproduction projects

Reality check: works vs. education

In Chinese LLM job markets, the weight of projects and portfolio depth is rising every year. In 2023, "knowing how to call APIs" got you past initial screening; by 2025, HRs want to see runnable repos, evaluation records, and verifiable numbers. Education determines whether you get in the door; your portfolio determines how high you stand once you're inside.

10. How to Use This List: An Action Plan ​

Reading JDs isn't about "having seen them" — it's about producing an executable plan:

PhaseActionOutput
Week 1Define 2–3 target role types (cross-reference with Module Overview); save 10 JDs for eachTarget role list
Week 1Circle ●●● skill words for each role type using the keyword radarHard skill list
Week 2Cross-reference with the Knowledge Breakdown, marking each as know/shaky/unknownPersonal study plan
Week 2Rewrite resume project titles and keywords to align with JD List requirementsAligned resume
Weeks 3–4Validate with Interview Questions; simulate one interview per weekInterview readiness baseline

11. Further Reading ​

Continue within the site

References ​

  • BOSS Zhipin — China's leading job platform; search "large language model/LLM/AI applications" for real-time JDs
  • Zhaopin — domestic recruitment platform with industry heat reports
  • OpenAI Careers — OpenAI's official career page; real JDs for overseas model roles
  • Anthropic Careers — Anthropic's official career page; includes Applied AI and other roles
  • Google Careers — Google's official careers page; Research Scientist / ML Engineer roles