Theme
JD List: LLM-Related Roles
Remember this one-liner from this page: a JD is the cheapest available sample of "what the market wants" — reading a hundred JDs is more reliable than guessing ten times what an interviewer will test on. This page organizes JD templates by role type, extracting the common skeletons of responsibilities and requirements that repeatedly appeared in major companies' public JDs since 2023, and includes a "keyword radar" so you can see at a glance the skill weight for each role.
Important Disclaimer
All JDs on this page are template frameworks distilled from common patterns in public major-company JDs — they do not point to any specific company's open positions, and do not represent the actual JD of any particular company. Real JDs vary widely in responsibilities, requirements, years of experience, and education thresholds depending on the company, team, and level. Always check each company's official career page for the latest, most accurate information (data current as of dataAsOf: 2025-08).
1. How to Use This List
1. JDs Change — Skill Combinations Are the Norm
JDs in the LLM field evolve faster than in any other industry: in 2023, "proficient with ChatGPT API" was the hot phrase; in 2024 it became "familiar with RAG and vector databases"; by 2025, "Agent and evaluation experience" was added on top. Focus on the unchanging fundamentals in JDs (Transformer principles, Python, engineering skills), then focus on the shifting combinations (RAG, Agent, multimodal, quantization). Fundamentals determine whether you can survive across cycles; combinations determine what you're worth right now.
2. Distinguish "Hard Requirements" from "Nice-to-Haves"
Major companies' JDs follow a pattern in their wording:
| Wording | Meaning | How to Respond |
|---|---|---|
| Proficient / Deep understanding / Skilled at | Hard requirement; interviewers will likely dig deep | Prepare a 30-minute deep-dive explanation |
| Familiar with / Understand | Soft requirement; explaining the concept is enough | Prepare a conceptual answer + one example |
| Experience with XX preferred / Bonus | Preferred but not required | Align with it in your resume and provide evidence |
| Master's degree or above / Top conference papers preferred | Education and academic thresholds | Break through with projects and portfolio if unmet |
3. Three Actions When Reading JDs
- Circle skill keywords: Pull out every technical noun and cross-reference them with the Knowledge Breakdown;
- Level them: Mark each skill as "I know it / I'm shaky on it / I don't know it," then generate a study list;
- Store samples: Save 3–5 JDs for your target role each week. After a month, you'll see the trend and know which way the market is moving.
2. Generic JD Skeleton: Four Blocks Every LLM Role Shares
Regardless of the role, major companies' JDs almost always have this structure:
| Block | Common Content | Deeper Intent |
|---|---|---|
| Responsibilities | Design/develop/optimize XX systems, keep up with cutting-edge tech, drive business adoption | Interview deep-dive area: What exactly did you do? What pitfalls did you hit? |
| Technical Requirements | Python/PyTorch, Transformer, fine-tuning, RAG, evaluation, deployment | Interview fundamentals area: maps to concepts and practice chapters on this site |
| Engineering Requirements | Code quality, distributed systems/serving, data processing, production stability | Tests "can you ship," not just "do you know concepts" |
| Soft Skills | Self-driven, communication, cross-team collaboration, English literature reading | Usually not strictly enforced, but project deep-dives will expose gaps |
Treat a JD as a "mini exam syllabus"
A JD is essentially a mini exam outline: responsibilities = interview deep-dive questions, technical requirements = fundamentals/definitions, engineering requirements = system design questions. Reverse-engineering interview questions from JDs yields much higher hit rates than randomly reading interview experiences.
3. JD Templates by Role Type
Each category below gives: role definition → responsibilities → requirements (required / bonus). All are common-pattern templates.
1. Algorithm Engineer · LLM Direction
Definition: Select and compose LLM capabilities (prompts, RAG, fine-tuning) for business scenarios, and handle evaluation and iteration.
Responsibilities:
- Design and deploy LLM solutions for business scenarios: scenario assessment, base model selection, prompt/fine-tuning/RAG strategy composition
- Build business evaluation sets, establish offline and online evaluation systems, and quantify production impact
- Keep up with the latest models and papers, assess their business applicability, and drive adoption
- Collaborate with engineering teams to complete model serving and ensure production stability
Requirements:
- Required: Proficient in Python; deep understanding of Transformer principles and fine-tuning methods (SFT/LoRA); familiarity with evaluation & benchmarks methodology; has delivered complete projects
- Bonus: Has RAG or Agent production experience; familiar with alignment (RLHF/DPO); has online A/B testing experience
2. NLP Algorithm Engineer
Definition: Develop text understanding and generation capabilities, using both classical NLP and LLM methods.
Responsibilities:
- Research and develop NLP capabilities: tokenization, entity recognition, text classification, semantic search, dialogue systems
- Integrate LLMs (few-shot, fine-tuning, RAG) into existing NLP pipelines and assess impact
- Build Chinese corpus processing and quality assessment pipelines
- Participate in model evaluation and online performance monitoring
Requirements:
- Required: Python and deep learning frameworks; familiarity with NLP fundamentals (tokenization, word embeddings, sequence labeling); understanding of language modeling and Transformers; has text-related project experience
- Bonus: Familiar with Chinese pre-trained models and corpus characteristics; has LLM application (RAG/fine-tuning) experience; familiar with vector search
3. LLM Training Engineer
Definition: Own the training systems and data pipelines for pretraining/continued pretraining/post-training (SFT, RLHF/DPO).
Responsibilities:
- Build training data pipelines: collection, cleaning, deduplication, mixing, quality monitoring
- Configure and tune training frameworks and distributed parallelism (data/tensor/pipeline/expert parallel)
- Monitor training dynamics (loss, gradients, throughput), diagnose and resolve non-convergence, overflow, and checkpoint recovery issues
- Keep up with scaling laws and latest training methods, evaluate and experimentally validate them
Requirements:
- Required: Solid deep learning fundamentals and training experience; familiar with PyTorch and distributed training (DeepSpeed/Megatron-LM); understand the full pretraining data pipeline; can independently troubleshoot training issues
- Bonus: Has MoE training or large-model experience; familiar with CUDA and memory optimization; has CUDA/HPC background; published top-tier conference papers
4. Inference Optimization Engineer
Definition: Optimize LLM inference performance and serving, improving throughput while reducing latency and cost.
Responsibilities:
- Research and deploy compression methods: model quantization (INT8/INT4), distillation, sparsification
- Optimize inference frameworks (vLLM / TensorRT-LLM / SGLang, etc.) for operators, batching, and KV cache management
- Conduct memory and performance analysis, stress testing, and tuning; build performance monitoring systems
- Address business-side inference cost and latency optimization needs
Requirements:
- Required: Deep understanding of the full Transformer inference pipeline (including KV cache and sampling); familiar with GPU architecture and CUDA programming; has performance analysis and tuning experience; proficient in Python and C++
- Bonus: Familiar with mechanisms like PagedAttention and continuous batching; hands-on quantization experience (GPTQ/AWQ); familiar with the full deployment & serving chain
5. AI Application Engineer (LLM Applications)
Definition: Integrate general-purpose LLMs into business, handling the production of prompts, RAG, Agents, fine-tuning, and evaluation.
Responsibilities:
- Develop LLM-powered business features: prompt design, RAG pipelines, tool calling
- Build business evaluation sets and regression tests, quantifying the impact differences between models/prompt/retrieval strategies
- Evaluate open-source models and frameworks, comparing the cost-effectiveness of self-deployment vs. API usage
- Optimize end-to-end latency and cost; ensure production stability
Requirements:
- Required: Proficient in Python; familiar with LLM APIs and open-source model deployment; understand the full RAG pipeline (indexing/retrieval/generation); has evaluation awareness and practice
- Bonus: Has Agent development experience; familiar with vector databases and embeddings; has fine-tuning (LoRA) hands-on experience; familiar with engineering practices (CI, monitoring, observability)
6. Agent Engineer
Definition: Design and develop LLM-based agents: tool calling, planning, memory, multi-agent collaboration.
Responsibilities:
- Design and implement Agent architectures: tool/function calling, task planning loops (ReAct, etc.), memory management
- Build Agent reliability: error recovery, guardrails, cost control
- Research and deploy Agent protocols and frameworks (function calling, MCP, etc.)
- Design protection mechanisms against prompt injection and other security risks
Requirements:
- Required: Python and systems engineering skills; understand LLM inference and prompt mechanisms; has experience with tool calling or workflow development; has complete Agent or automation product experience
- Bonus: Has memory (long-term/short-term) and multi-agent design experience; has production stability experience with LLM-based Agents; familiar with vector search and evaluation
7. Evaluation Engineer
Definition: Build evaluation systems for model and product quality, providing the basis for model selection and iteration.
Responsibilities:
- Build evaluation systems: public benchmarks + custom business sets + human/model judgment workflows
- Run offline batch evaluations and regression tests to prevent degradation from model iterations
- Design and analyze online evaluations (A/B testing, user feedback), quantifying production impact and risks
- Address benchmark contamination and evaluation bias issues; maintain evaluation data quality
Requirements:
- Required: Understanding of evaluation & benchmarks methodology; familiar with common benchmarks (MMLU/GSM8K/HumanEval, etc.) and evaluation tools (lm-eval-harness/OpenCompass); Python and data analysis skills
- Bonus: Has LLM-as-a-judge production experience; has data labeling workflow management experience; familiar with hallucination and safety evaluation
8. Data Engineer · LLM Direction
Definition: Handle collection, cleaning, mixing, and quality assurance of pretraining/post-training corpora and business data.
Responsibilities:
- Build large-scale corpus pipelines: collection, parsing, cleaning, deduplication, filtering, mixing
- Produce instruction/preference data and ensure quality (including managing labeling teams and synthetic data)
- Build data quality monitoring and compliance review mechanisms
- Support the model team's data needs by providing reusable data infrastructure
Requirements:
- Required: Python and data processing skills (Spark/Flink, etc.); familiar with deduplication (MinHash) and filtering methods; understand pretraining data principles; has data pipeline engineering experience
- Bonus: Has synthetic data experience; understands copyright and compliance risks; has search/recommendation data processing experience
4. Keyword Radar: High-Frequency Skill Words by Role
The "frequency" ratings in the radar table are subjective assessments based on how often each skill appears across JD samples (●●● = high-frequency hard requirement / ●● = common / ● = bonus or team-dependent). Sample observations come from public major-company JDs between 2023–2025, and are for exam-preparation direction only.
| Role | High-Frequency (●●●) | Common (●●) | Bonus (●) |
|---|---|---|---|
| Algorithm Engineer (LLM) | Transformer, fine-tuning/LoRA, RAG, evaluation, Python/PyTorch | Prompt engineering, Agent, alignment | Multimodal, papers, A/B testing |
| NLP Algorithm Engineer | NLP fundamentals, Transformer, text classification/extraction, Python | Vector search, fine-tuning, evaluation | Dialogue systems, knowledge graphs |
| LLM Training Engineer | Distributed training, pretraining, data pipelines, PyTorch/DeepSpeed | Parallelism strategies, loss diagnostics, MoE | CUDA, papers, supercomputing experience |
| Inference Optimization Engineer | Quantization, KV cache, vLLM/TensorRT-LLM, performance analysis, C++/CUDA | Continuous batching, operator optimization, GPU architecture | Inference framework source contributions |
| AI Application Engineer | LLM API, RAG, prompt engineering, Python, vector databases | Agent, LoRA fine-tuning, evaluation | Deployment ops, CI/CD |
| Agent Engineer | Tool calling, ReAct, memory, Python, systems engineering | MCP, multi-agent, evaluation | Security/injection defense, reinforcement learning |
| Evaluation Engineer | Benchmark sets, evaluation methodology, lm-eval-harness, data analysis | LLM-as-a-judge, regression testing | Manual labeling management, online A/B |
| Data Engineer (LLM) | Data processing, cleaning & dedup, Spark/Flink, data pipelines | MinHash, filtering, compliance | Synthetic data, labeling management |
How to use the radar table
Treat the "●●●" column as mandatory interview territory, the "●●" column as "at least understand the principles," and the "●" column as "bonus on your resume if you have it, no loss if you don't." Cross-reference with the Knowledge Breakdown and check each item off.
5. Master List of High-Frequency Skill Words
Looking at all roles together, here are the highest-frequency skill words (coverage = number of role categories where it appears as a required skill):
| Skill Word | Coverage | Related Page on This Site | Interview Weight |
|---|---|---|---|
| Python | 8/8 | — (basic skill) | Mandatory for coding rounds |
| Transformer/Attention | 7/8 | Transformer Architecture Explained | Core mandatory |
| Fine-tuning (SFT/LoRA) | 6/8 | Fine-tuning / Fine-tuning Practice | High-frequency |
| Evaluation | 6/8 | Evaluation & Benchmarks / Evaluation Practice | High-frequency, easy to stand out |
| RAG | 5/8 | RAG Practice | High-frequency |
| Prompt Engineering | 5/8 | Prompt Engineering | High-frequency |
| Distributed Training | 3/8 | Framework & Tool Selection | Mandatory for training roles |
| KV cache / Quantization / Deployment | 3/8 | Deployment & Serving | Mandatory for inference roles |
| Agent / Tool Calling | 3/8 | LLM-based Agents | Fastest growing |
| Alignment (RLHF/DPO) | 3/8 | Alignment | Deep-dive topic |
6. Reading Between the Lines: Filler vs. Substance in JDs
Not every sentence in a JD carries the same credibility. Here's how to distinguish common "filler" from "substance" in major-company JDs:
| JD Phrasing | Real Meaning | Priority |
|---|---|---|
| "Good teamwork and communication skills" | Generic boilerplate — everyone writes this | Ignore |
| "Strong self-drive and learning ability" | Tech stack may not be defined yet; you'll be figuring things out on the job | Reference only |
| "Understanding of mainstream LLMs and their applications" | You've at least used them and understand basic principles | Substance — prepare for principle-level questions |
| "Experience with LLM applications/fine-tuning/RAG projects" | Need project evidence; interviewers will dig deep | Core substance |
| "Familiar with distributed training frameworks" | Hard threshold for training roles; mandatory fundamentals | Core substance |
| "Top conference publications / open-source contributions preferred" | Differentiator; nice to have but not required | Bonus |
| "Can handle pressure / fast iteration pace" | Signal for overtime and version rush deadlines | Warning flag |
Three signals that a JD means "we're seriously hiring":
- Technical requirements are specific (naming specific frameworks, specific tasks) → the team knows what they want
- Clear deliverable descriptions ("responsible for XX system," "improve XX metric") → there's a real opening
- "Required" and "bonus" items are clearly tiered → they have mature talent standards
Conversely, JDs that are all vague boilerplate and loaded with "strong learning ability" are likely HR mass-postings — research the team's actual situation before applying, and don't waste time on interviews for positions that are probably just backup candidates.
7. Trend Observations (2023–2025)
- "Knowing how to use" is depreciating; "can prove effectiveness" is appreciating: Early JDs said "familiar with ChatGPT," now they say "capable of evaluation and regression testing." Interviews increasingly ask "how do you prove your solution works."
- RAG has moved from bonus to mandatory: Almost all application roles now require RAG production experience; "Agent" and "tool calling" requirements have been added since 2024.
- Training roles are consolidating at top companies: Pretraining-only roles exist only at a handful of model companies; most JDs have shifted to "fine-tuning/post-training on open-source models."
- Inference optimization has become its own role: Keywords like quantization, KV cache, and vLLM have risen dramatically in frequency over two years; candidates with systems backgrounds have grown more competitive.
- Data and evaluation roles have gone from "afterthoughts" to "first-class citizens": Increasingly many teams write data quality and evaluation systems into their JDs with formal leveling.
8. Quick Reference: English JD Keywords
For roles at foreign companies and global-facing positions, here are the most common keywords in English JDs (with Chinese equivalents and interview tips that map directly to testable topics on this site):
| English Keyword | Meaning | Chinese Equivalent | Interview Tip |
|---|---|---|---|
| LLM / Foundation Model | Large language model / foundation model | Large language model | Transformer principles are mandatory |
| Prompt Engineering | Prompt design | Prompt engineering | Often paired with in-context learning |
| Retrieval-Augmented Generation | RAG | Retrieval-augmented generation | Full pipeline is mandatory |
| Fine-tuning / PEFT | Fine-tuning / parameter-efficient fine-tuning | Fine-tuning | LoRA principles are mandatory |
| LoRA / QLoRA | Low-rank adaptation and its quantized variant | LoRA | Rank, alpha, memory comparison |
| RLHF | Reinforcement learning from human feedback | Alignment | Three-step process is mandatory |
| DPO | Direct preference optimization | Alignment | Compare with RLHF |
| Agent / Tool Use / Function Calling | Agent / tool calling | Agent | ReAct loop and reliability |
| Evals / Benchmark | Evaluation / benchmark | Evaluation | Contamination and LLM-as-a-judge are bonus points |
| Context Window / Long Context | Context window / long context | Long context | RoPE extrapolation and RAG tradeoffs |
| Quantization | Quantization | Inference optimization | INT8/INT4 tradeoffs |
| Inference Optimization | Inference optimization | Inference optimization | KV cache, batching, throughput |
| Latency / Throughput | Latency / throughput | Deployment metrics | TTFT/TPOT |
| Vector Database / Embedding | Vector DB / embeddings | RAG infrastructure | Selection logic |
| Distributed Training | Distributed training | Training engineering | DP/TP/PP/ZeRO |
| Hallucination / Grounding | Hallucination / grounding | Hallucination | Causes and mitigation |
| Production / MLOps | Production environment / engineering | Engineering capability | Evaluation and observability loops |
Reading between the lines of English JDs
"Strong CS fundamentals" = you'll face tough coding challenges; "Experience shipping to production" = deep-dive into production stability and incident handling; "Familiar with research literature" = they'll test your paper reading and reproducibility skills. English JDs use subtler language — prepare for the testable topics behind each keyword.
9. Hard Thresholds: Education, Experience, and Equivalent Work
Thresholds are statistical patterns, not hard rules — but knowing the patterns helps you allocate your limited resume space wisely:
| Stage | Common Thresholds (common patterns across major-company public JDs) | What to Do If Unmet |
|---|---|---|
| Fresh graduate | Master's is mainstream; algorithm/training roles generally require Master's or above; application/engineering roles accept Bachelor's | Prove hands-on ability with runnable, complete projects + evaluation data |
| 1–3 years | "3+ years LLM/NLP/recommendation/search experience" appears frequently | Deep projects can substitute for years: one thoroughly explained project ≈ one year of experience |
| 3+ years / Senior | "Independently responsible for complete systems," "experience with large-scale production" | Focus interviews on production stability and post-incident retrospectives |
| PhD / Top conference | Hard threshold or strong bonus for training and research roles | Fill paper gaps with open-source contributions and reproduction projects |
Reality check: works vs. education
In Chinese LLM job markets, the weight of projects and portfolio depth is rising every year. In 2023, "knowing how to call APIs" got you past initial screening; by 2025, HRs want to see runnable repos, evaluation records, and verifiable numbers. Education determines whether you get in the door; your portfolio determines how high you stand once you're inside.
10. How to Use This List: An Action Plan
Reading JDs isn't about "having seen them" — it's about producing an executable plan:
| Phase | Action | Output |
|---|---|---|
| Week 1 | Define 2–3 target role types (cross-reference with Module Overview); save 10 JDs for each | Target role list |
| Week 1 | Circle ●●● skill words for each role type using the keyword radar | Hard skill list |
| Week 2 | Cross-reference with the Knowledge Breakdown, marking each as know/shaky/unknown | Personal study plan |
| Week 2 | Rewrite resume project titles and keywords to align with JD List requirements | Aligned resume |
| Weeks 3–4 | Validate with Interview Questions; simulate one interview per week | Interview readiness baseline |
11. Further Reading
Continue within the site
- Knowledge Breakdown — turn each keyword in the radar into a personalized study plan
- Resume Analysis — rewrite your experience into evidence aligned with JD requirements
- Interview Questions — the ultimate test of whether the JD requirements are met
- Module Overview & Career Landscape — how to choose between roles; full picture of salaries and trends
References
- BOSS Zhipin — China's leading job platform; search "large language model/LLM/AI applications" for real-time JDs
- Zhaopin — domestic recruitment platform with industry heat reports
- OpenAI Careers — OpenAI's official career page; real JDs for overseas model roles
- Anthropic Careers — Anthropic's official career page; includes Applied AI and other roles
- Google Careers — Google's official careers page; Research Scientist / ML Engineer roles