Appearance
Reading Discipline & FAQ
One-sentence positioning: the most important thing about reading inference papers is not "how many you read" but "whether you have a repeatable reading discipline" — knowing which speedups to trust, which profile figures to see through, and which papers to throw away. This page first sets the discipline, then answers your high-frequency questions from "can't understand" to "can't finish."
How to use this page
Read the discipline part (Section 1) closely once, then consult the "red-flag table" every time you read a paper; the FAQ part does not need to be read start to finish — come back when you hit the matching problem. Treat it as a methodology handbook, not an article you must finish in one sitting.
1. Reading Discipline: The Judge of Speedups, Profiles, and Papers
1.1 How to Read a Speedup: Every Speedup Is a Conspiracy of "Model + Batch + Hardware + Protocol"
The most eye-catching things in inference papers are the numbers: "2-3x faster than vLLM," "6.5x speedup," "4x throughput." But remember one iron law:
A speedup is never an absolute value of "how good the method is" — it is the joint product of five things: "method + model + batch + hardware + evaluation protocol." Change any one condition and the number is no longer comparable.
So, facing any speedup, force yourself to ask four questions:
| Question | What it means | Why it matters |
|---|---|---|
| Model | Which model? How many parameters? How long a context? | A 70B and a 7B have completely different bottlenecks — 70B is memory-bound, 7B is compute-bound |
| Batch | What batch size? Single user or many? | The optimal algorithm at batch=1 vs. batch=256 is completely different — the former is decode optimization, the latter prefill optimization |
| Hardware | A100/H100/B200? Which generation? | FA2 is optimal on A100, FA3 on H100, and FP4 is only sweet on B200 |
| Protocol | How is the metric computed? Single run or averaged over many? Is the standard deviation reported? | Reporting only the best run without variance may just be luck |
Pay special attention to relative vs. absolute improvement: when a paper says "30% improvement on X," it might be from 100 ms to 70 ms (a significant absolute change) or from 1 ms to 0.7 ms (almost meaningless in absolute terms). Always convert to absolute numbers before judging.
Another numerical trap: kernel micro-benchmark vs. end-to-end benchmark. The "2-4x" reported in the FlashAttention paper is end-to-end training speedup; many follow-ups report only kernel micro-benchmarks (e.g. "our attention kernel is 1.3x faster than FA2"), but end-to-end it may be only 1.05x — because attention is only 30% of the total time, so a 1.3x kernel gives only 1.05x end-to-end. Always ask "is this speedup kernel-level or end-to-end."
A one-sentence mnemonic
Take away the mechanism, not the number. The number is the product of "that time, that place, that machine"; the mechanism (why it works, where its boundary is) is what holds across time and space. This is the reading view emphasized throughout the Start Here section.
1.2 How to Read a Profile Figure: Axes First, Then Trend, Then Conclusion
A systems paper's figures are denser in information than an algorithms paper's: one roofline figure, one batch-throughput curve, or one latency CDF can directly expose the method's true level. Read in three steps:
text
Step 1: read the axes Is the x-axis log or linear? Where does the y-axis start and end? What are the units?
(a log axis "flattens" differences; a linear axis amplifies early jitter)
A compute axis or a time axis? A latency axis or a throughput axis?
Step 2: read the trend Is the curve monotonically rising or U-shaped after overfitting?
Do the methods "pull apart" or "run neck and neck"?
Does it still hold at long sequences / large batches?
Step 3: read the conclusion What conclusion does the author draw from this figure?
Does the figure really support it? Is there a more honest way to plot it?Three frequent "figure-reading" traps:
- Reporting only the rising curve, not the absolute values: in a figure with batch size on the x-axis and throughput on the y-axis, if the y-axis baseline is truncated, the curve looks much steeper. Watch the axis ranges.
- In ablation figures, look at "which removal hurts most": the point of an ablation is to answer "how much does each component contribute." The reading is not "how strong the full model is" but which component, when removed, drops accuracy the most — that is the method's true core innovation. Components whose removal barely matters are often icing on the cake, or even dispensable.
- In a roofline figure, look at arithmetic intensity: the goal of inference optimization is to push a method from "memory-limited" to "compute-limited." Reading a roofline figure, first check whether the method lands left of the knee (IO-bound) or right (compute-bound), and whether the optimization crosses the knee. An optimization that does not cross the knee has a speedup ceiling equal to the bandwidth ratio — the hard ceiling roofline gives. See The Roofline Model and Compute Analysis and The GPU Memory Hierarchy and the Bandwidth Wall.
1.3 How to Judge a Paper's Quality: The Red-Flag Table
Read enough inference papers and you will find that good and bad papers diverge clearly in "degree of honesty." The table below lists signals worth guarding against — the more that appear, the more you should lower your trust level:
| Red flag | Manifestation | Response |
|---|---|---|
| Speedup without protocol | Claims "3x speedup" without stating model, batch, hardware, or evaluation script | Check the code/appendix directly; if unfindable, mark "not reproducible" |
| Weak baseline | Compares only "an old unoptimized method," avoiding the strongest contemporary method | Cross-check against the latest vLLM/SGLang/TensorRT-LLM benchmarks |
| No ablation | The method stacks five or six modules but never does "what if we remove one" | The core contribution cannot be located; the novelty is in doubt |
| Single model/hardware | Reports results only on Llama-70B (or one model), only on A100 | High chance of overfitting to one model / one hardware generation |
| No variance reported | Every experiment reports "the best run," with no mean and no standard deviation | Gains from a lucky random seed are not trustworthy |
| No limitations discussed | The paper has no Discussion/Limitation section | Either not thought through, or hiding problems |
| Relative numbers in the abstract | "30% improvement," "substantially outperforms," but no absolute numbers in the body | Likely a tiny base or an unfair comparison |
| Code perpetually "coming soon" | The paper says "code will be released" but no repo exists | Reproduction is the minimum bar; be extra wary without code |
| Kernel-level speedup but no end-to-end | Reports only "my attention kernel is 1.3x faster," not end-to-end | Likely only 1.05x end-to-end, because attention is 30% of the total |
| Mixing batch=1 with large batch | Implies a batch=1 speedup also holds at large batch | The optimal algorithms for decode (batch=1) and prefill (large batch) are completely different |
Red flags != the paper is bad
Red flags are indicators that "you need to raise your guard," not a verdict to "shoot it down on sight." A paper with weaknesses can still contribute a valuable idea — its method may fail in other settings, but the problem definition and the thinking may be right. The correct posture: note the red flags, and after finishing, label the paper with a "credibility rating" — not black-and-white.
2. FAQ: Twelve Questions on the Road of Reading Papers
Q1. How much CUDA and systems background do I need before reading inference papers?
Less than you imagine, but more than for reading algorithms papers. Reading an inference paper requires three kinds of foundation:
- The CUDA programming model: the hierarchy of thread / warp / thread block / SM, the difference between shared memory and global memory, warp synchronization and reduction. These are in the first 5 chapters of the CUDA Programming Guide. You do not need to write CUDA kernels, but you must be able to read the "thread-block partitioning diagrams" in papers.
- The GPU performance model: roofline, arithmetic intensity, the HBM-SRAM bandwidth gap (~30x), Tensor Core compute and precision. These are in GPU Architecture and Optimization and The Roofline Model and Compute Analysis.
- Systems fundamentals: virtual memory, paging, batching, scheduling. vLLM's PagedAttention borrows OS virtual memory wholesale; Orca's iteration-level scheduling borrows OS process scheduling wholesale. Understand the basic OS concepts and you understand half of systems papers.
The real trick is to consume the math and the engineering in layers: Pass 1 reads only the conclusions, not the pseudocode; Pass 2 captures "the loop order in Algorithm 1"; only when you plan to reproduce or improve the method do you need to pore over the CUDA kernel line by line. The algorithm layers of most entry papers (FA1, vLLM, GPTQ) are not complicated — the engineering is.
Q2. What if I can't understand it?
First accept a fact: not understanding is the default state, not an anomaly. Inference papers are not textbooks; authors assume readers are fellow systems researchers and skip large amounts of setup. The correct processing order is:
(1) Switch materials — the same paper usually has blog explainers, video walkthroughs, and illustrated notes; use second-hand materials to build the global map first. Tri Dao's blog, the vLLM official blog, the FlashInfer blog, and the SGLang blog are all excellent crutches.
(2) Go back and fill prerequisites — not understanding is often not this paper's fault but a missing prerequisite. Reading FlashAttention and stuck on online softmax? Go read Milakov & Gimelshein 2018, "Numerically Stable Softmax." Reading vLLM and stuck on PagedAttention? Go review the OS virtual-memory chapter.
(3) Mark, skip, and come back — note what you do not understand, finish the parts you do, and many questions dissolve on a second reading. Reading the same paper three times beats reading three papers once.
A "switch materials" list for when you are stuck
- FlashAttention: Tri Dao's blog, illustrated explainers such as the Stat-Ops writeups
- vLLM / PagedAttention: the vLLM official blog, Lilian Weng's PagedAttention explainer
- Speculative decoding: Lilian Weng's decoding survey, the SGLang blog
- GPTQ/AWQ: the HuggingFace quantization blog, EleutherAI's quantization writeups
Q3. Should I read the original English?
Yes, and the sooner the better. Three reasons: (1) translation loses information — terms like "warp-specialized," "online softmax," and "chunked prefill" have no standard Chinese rendering, and Chinese explainers often lag by months or do not exist; (2) inference deployment has no reliable complete-Chinese-translation ecosystem — frontier papers always arrive in English first; (3) reading the original is "pay once, benefit for a long time" — after 20 English systems papers, your reading speed jumps qualitatively, because systems-paper English is highly formulaic ("we propose...", "we observe that...", "to the best of our knowledge...").
The advice for beginners is "mix Chinese and English": read a Chinese blog first for the global map, then return to the English original to check the details, treating the original as the single source of truth. Pay special attention to NVIDIA's official blogs and white papers — they are the best crutches for hardware-co-design papers on H100/B200/FP8/FP4.
Q4. How do I take notes?
The goal of notes is "still understandable three months later," not "felt clear at the time." A structured "one-page note" template:
text
Paper title / year / venue
One sentence: what bottleneck does it solve? (in your own words)
Core mechanism: what is the method's core idea? (1-3 sentences + a diagram you drew yourself)
Key numbers: the 2-3 most important speedups + model/batch/hardware/protocol
Ablation conclusion: which component contributes the most?
Boundary of applicability: under what batch, sequence length, and hardware generation does it hold?
Limitations: what the authors admit + what you found
Relevance to me: what is it good for in my deployment/research?
Credibility rating: high / medium / low + whyStrongly recommended: draw the mechanism diagram by hand — use arrows and boxes to draw the paper's method; if you can draw it, you truly understand it. FlashAttention's two-level loop, vLLM's block table, and EAGLE's draft tree should each be drawn once. The note tool does not matter — Obsidian, Notion, or plain Markdown all work; the key is a fixed template per paper so notes are batch-searchable. Every entry in Classic Papers in Depth follows the fixed structure "problem background -> core mechanism -> experimental conclusions -> implications," which you can use directly as your note template.
Q5. How do I find related papers?
Use the "start from one paper, expand in both directions" method — far more efficient than random searching:
text
The paper for method X
+- backward: look at its references (References) --> find earlier source papers
+- forward: look at who cites it (Citations) --> find later improvement workConcrete tools: Google Scholar for citations and related articles; Semantic Scholar has a good API and recommendations; Connected Papers shows a paper's citation network as a graph, good for quickly locating "the core cluster of papers" in a direction. Papers with Code is the best entry for "task -> SOTA -> code -> paper" — follow a benchmark's leaderboard to find all major methods for that task. Pair it with this site's Paper Map to build a global coordinate system first, then let citation chains take you deep.
Inference deployment has two more distinctive channels: (1) GitHub issues/PRs — the issue sections of vLLM, SGLang, and TensorRT-LLM are full of grounded engineering discussions like "how to do Y with X," often more down-to-earth than papers; (2) official blogs — the blogs of Tri Dao, Lilian Weng, HuggingFace, and SGLang are often more readable than the papers themselves and are excellent second-hand materials.
Q6. How do I judge whether a paper is outdated?
Three dimensions: (1) citations and follow-ups — check its citation count and who cites it on Semantic Scholar or Google Scholar; if later work has clearly overturned or greatly improved it, downgrade it. FlashAttention v1 is still cited by FA2/FA3 and counts as "classic, not outdated"; Orca was extended by vLLM and others in 2023 — "integrated, but the idea still matters." (2) hardware and compute context — FlashAttention v1 is still optimal on A100 but has been replaced by FA3 on H100; GPTQ is still mainstream on A100 but often replaced by FP8 on H100. Watch whether the paper's "hardware premise" still holds. (3) surveys and benchmarks — if the latest survey on the task no longer mentions the paper, or the latest SOTA on a benchmark uses a completely different paradigm (e.g. from PagedAttention to vAttention), it has entered the "historical value" stage. The point of judging obsolescence is not "don't read old papers" but "give them the right positioning" — old papers are often the key to understanding new ones; Classic Papers in Depth and Frontier Advances are organized on this "past-present contrast" idea.
Q7. Should I reproduce the experiments?
Depends on your goal — three intensities:
| Goal | Reproduction intensity | Notes |
|---|---|---|
| Entry-level learning | No reproduction needed | Understand + notes + explain to someone; spend your time on breadth |
| Engineering deployment | Must reproduce (at small scale) | Without running it once, you never know how many "hidden details" (CUDA kernel tuning, block-size choice, calibration data) the paper left unwritten |
| Research | Deep reproduction + ablations | The entry ticket of research: without reproduction, all follow-up improvements are castles in the air |
Key insight: the cost of one reproduction is often the true measure of the paper's value. If a method's "official code cannot reproduce the paper's numbers," that information matters more than the paper. Standing in 2024, prioritize papers with an official open-source implementation — use Curated Resources or Papers with Code to find official code, saving an order of magnitude of time over implementing from scratch.
For inference papers, "reproduction" has three levels: (1) run the official repo + official weights to reproduce the paper's numbers (shallowest); (2) run the official repo + your own model/hardware and see whether the numbers still hold (medium); (3) write a simplified kernel or scheduling snippet from scratch to verify whether the mechanism really works (deepest). Engine builders need level 3; application builders need only level 2.
Q8. What if there are too many papers to finish?
First break an obsession: "finishing all the papers" was never the goal, and not finishing does not mean failure. Papers are a "fetch on demand" resource, not a "must-clear checklist." Three practical rules:
(1) The 80/20 rule — 20% of papers deserve a close read, 80% need only "read the abstract + look at the profile figure + note one sentence"; a close read is a solid two hours, a skim ten minutes.
(2) Read by task, not by heat — set yourself a current topic (e.g. "prefill/decode disaggregation") and read only papers in that topic; write a summary before switching, far more effective than "randomly skimming two today."
(3) Tolerate "read and forget" — the brain is not a database; the key is to make notes your external memory. For concrete route planning see Reading Paths. Do not be greedy — never deeply read more than two or three papers at once.
Inference deployment has another distinctive phenomenon: much "important work" appears in GitHub PRs before papers. So when you "can't finish the papers," you can turn to "can't finish the PRs" — the main-branch PRs of vLLM/SGLang hide a wealth of engineering detail, often more grounded than papers.
Q9. How do I explain a paper to someone?
The Feynman technique, best practiced on paper reading. Before explaining, ask yourself: can I state in three sentences "what bottleneck this paper solves, how it solves it, and how well it works"? If not, you do not understand it yet. Concrete steps:
- Explain to a 5-year-old (or your non-technical friend): only the motivation and intuition, no jargon — forcing you to find the method's core intuition in the plainest language. Example: FlashAttention = "keep the intermediate results in fast memory and make fewer trips to slow memory."
- Explain to a peer (colleague / classmate / interviewer): the mechanism, experiments, and limits — forcing you to think through every design choice and its alternatives. Example: FlashAttention's "why is the outer loop over K/V blocks rather than Q blocks."
- Write a 300-word summary: speaking is easier to bluff than writing; only by writing do you find where you are stuck.
Especially practical for interviews: in inference-deployment interviews, "talk through a systems paper" is almost guaranteed, and the standard answer structure is "problem -> intuition -> mechanism -> experiments -> limitations -> what you would do differently." If you cannot produce it, your note's "relevance to me" section was left blank.
Q10. How do I quickly judge whether an inference paper is worth reading?
Before a close read, run it through the "three-minute funnel":
text
Step 1 Title + abstract (30s) -> is the bottleneck relevant to me? Can the method be stated in one sentence?
Step 2 Profile figures + conclusion (60s) -> is the speedup significant? Under what model/batch/hardware?
Step 3 Last paragraph of the intro (30s) -> is the authors' contribution list concrete?
Step 4 Check code/reproduction (60s) -> is there official code? Star count? Has anyone reproduced it?If two or more of the four questions go unanswered or get a negative answer, put it back on the "to-read" list. Judging "not worth reading" is as important as judging "worth reading" — leave your limited time for the truly important papers. Do not be held hostage by the words "3x speedup"; the speedup leaderboard changes every month, and a paper's value is far more than its benchmark rank.
For inference papers, there are a few extra high-value signals:
- Authors from first-line systems groups: the Princeton/Dao group, the Berkeley/Lianmin Zheng group, the DeepSeek systems group, the NVIDIA TensorRT-LLM team, the vLLM core team — work from these groups is almost always worth reading.
- Integrated into a mainstream engine: work already in the main branch of vLLM, SGLang, or TensorRT-LLM has almost always been industrially validated.
- Ships with open weights: such as DeepSeek-V3, Mixtral — work with open weights tends to have more impact.
Q11. Should I read the survey or the original?
Survey first, then originals — the survey is the map, the originals are the terrain. The value of a survey: (1) a few pages give you a direction's complete coordinate system — who did it first, who improved what, and where it is stuck now; (2) it builds your terminology so unfamiliar concepts no longer trip you when reading originals; (3) a survey's reference list is itself a high-quality reading list.
But a survey's weakness is insufficient depth and possible lag — after 2024, inference systems change so fast that a survey can be outdated before it is finished. So the correct posture is: for a new field, read a survey from the last two years to build the map, then immediately switch to the two or three originals you care about.
To find surveys, search arXiv for "survey" + topic (e.g. "speculative decoding survey," "LLM quantization survey"), or follow the surveys mentioned in Frontier Advances. This site's Paper Map can also serve as a condensed survey — it arranges a decade of key papers by theme into one table.
Q12. How many papers count as "entry-level"?
There is no fixed number, but there are verifiable milestones. A practical criterion: when you can, without any material, sketch the "bottleneck -> key methods -> mainstream routes -> current bottleneck" structure of a subfield and point out three or more representative works and their relationships, that subfield counts as entered.
By this site's suggestion, the entry path is roughly:
- The ten classics (FlashAttention to DeepSeek-V3) build the timeline ->
- Pick one direction you are interested in and closely read 5-10 papers ->
- Follow that direction's latest work from the past year.
Refer to the arrangement of Classic Papers in Depth and the Paper Map; a single direction usually needs 15-30 high-quality papers to support "can follow academic discussions and talk through them in an interview." The bar for "entry" is "can hold a conversation," not "finished reading everything."
3. Papers vs. Engineering: Why "Paper Readers" Build Engines
A common confusion: "I do engineering deployment and don't write papers — what good is reading papers to me?" The answer: papers -> framework implementation -> deployment tuning are three stacked layers, and missing any one means you cannot go deep.
text
Paper layer: mechanism + experiments + limitations
| papers give the "why"
Framework layer: implementations of vLLM/SGLang/TensorRT-LLM
| frameworks give the "how"
Deployment layer: your model + your hardware + your load
| deployment gives the "what to tune"
You: tuning + selection + troubleshootingPeople who only do the deployment layer can only "trial and error + search issues" when they hit a bottleneck; people who have read papers can reason backward from the mechanism to where the bottleneck is, why this engine underperforms on this load, and whether to switch engines or write a kernel. This is why interviewers love to ask "which engine's underlying papers have you read" — it directly separates "someone who can use an engine" from "someone who can modify an engine."
See Inference Engine Comparison, Tuning and Performance Optimization, and Deploy an Inference Service from Scratch.
4. When to Chase the Frontier
Not every situation calls for chasing the frontier; here are strategies for different scenarios:
| Scenario | Chase the frontier? | Strategy |
|---|---|---|
| Firefighting / online incident | No | Use the tools and configs you know well; methods from new papers are unvalidated — do not try them while firefighting |
| New project selection | Moderately | Read surveys and mainstream-engine release notes from the last 6 months, avoiding options that "look new but are already outdated" |
| Writing a survey / tech talk | Must | Read at least all the important work of the last year, or the survey adds no information |
| Cramming before an interview | Moderately | Focus on the ten papers of Classic Papers in Depth; for the frontier, knowing the keywords suffices |
| Doing research / writing papers | Continuously | Fix a weekly slot for scanning arXiv + repo PRs, or your work may already have been done |
| New-hardware adaptation | Must | A hardware-generation shift means algorithm rewriting; when new hardware (H100->B200, A100->Ascend 910B) ships, you must track the new operator papers |
Do not chase the frontier while firefighting
When an online service is down, use the tools and configs you know best — not the new methods from papers. The reason: new methods are unvalidated on your load, your data, and your hardware, and may introduce unknown problems. Chasing the frontier is a long-term investment, not firefighting medicine.
5. Summary: Paper Reading Is Long-Termism
Viewed on the timescale, reading inference papers is clearly compounding:
text
Short term (weeks): read a few, fail to understand, doubt yourself -> everyone's rite of passage
Mid term (months): closely read 20 papers + notes + retelling -> noticeably keep up with discussions
Long term (1-2 years): 50-100 papers in one direction -> can judge directions and pose questions yourselfThe compound interest of paper reading shows in three places: speed — the first FlashAttention takes a week, the tenth attention variant takes an hour, as formulaic reading gets faster and faster; judgment — the more you read, the faster you spot weak baselines and flashy speedups, and the red-flag table becomes intuition; expression — the ability to explain a paper clearly is the scarce expression skill you use in interviews, reviews, and collaboration.
Back to the starting point of this FAQ:
Three disciplines — the three sentences most worth taking away
- Take away the mechanism, not the number — every speedup is a conspiracy of model + batch + hardware + protocol.
- Not understanding is the default state — switch materials, fill prerequisites, mark and reread; do not fight one paper to the death.
- Notes are external memory, and retelling is the only test — reading without notes is as good as not reading, and understanding you cannot articulate is as good as none.
Being unable to finish, unable to understand, or forgetting right after reading are not problems. There is only one problem: whether each session leaves behind a little something reusable. If yes, that is long-termism; if no, reading ten thousand papers is just writing on sand.
Further Reading
- Start Here — the entry to the papers section and the three onboarding routes
- Reading Paths — specific paper lists and reading orders arranged by goal
- Paper Map — a ten-year timeline coordinate system of key inference papers
- Classic Papers in Depth — close readings of the ten field-changing papers, doubling as a note template
- Frontier Advances — tracking the 2024-2026 frontier of FP8/FP4, prefill/decode disaggregation, EAGLE-3, and new hardware
- What Is Inference Acceleration? — the conceptual foundation of inference deployment
- Hardware Primer — background on GPUs, NPUs, and domestic hardware
- Glossary — a quick terminology lookup while reading
- Learning Paths: Three Routes — plan your long-term learning from the whole site's perspective
- Inference Engine Comparison — the bridge between papers and engineering: mainstream engine comparison
- Inference Benchmarking in Practice — how to run inference benchmarks
- Curated Resources — entry points for finding paper code, benchmarks, tools, and official code
References
The following are all real public resources for self-study:
- arXiv preprint server — the first publication venue for most inference papers; focus on
cs.LG/cs.DC/cs.OS - Google Scholar — general academic search for citation counts, cited-by, and related articles
- Semantic Scholar — semantic paper search and recommendation, with an API
- Connected Papers — explore a direction's paper landscape as a citation-network graph
- Papers with Code — papers + code + benchmark aggregation; cross-check SOTA and reproduction status
- CUDA Programming Guide — the hardware background required for reading operator papers
- NVIDIA developer blog — the best crutch for hardware-co-design papers on H100/B200/FP8/FP4
- Tri Dao's homepage — the FlashAttention author's blog
- vLLM project blog — official explainers of PagedAttention and later engineering
- SGLang blog — official explainers of RadixAttention and speculative decoding
- FlashInfer blog — official explainers of block-sparse and composable KV formats
- Lilian Weng's blog — systematic surveys of LLMs and inference optimization
- HuggingFace quantization blog — Chinese explainers of GPTQ/AWQ/SmoothQuant
- Keshav. How to Read a Paper (2007) — the original paper of the "three-pass method," required reading for methodology
- Dao et al. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (NeurIPS 2022)
- Kwon et al. Efficient Memory Management for LLMs Serving with PagedAttention (SOSP 2023)
- Frantar et al. GPTQ: Accurate Post-Training Quantization for GPT (ICLR 2023)
- Lin et al. AWQ: Activation-aware Weight Quantization (MLSys 2024)
- Li et al. EAGLE-2: Faster Speculative Decoding with Dynamic Draft Trees (EMNLP 2024)