Every major capability in today's AI stack — the transformer inside every chatbot, the
retrieval trick behind every RAG pipeline, the reasoning loop behind every coding agent — traces
back to a specific, citable research paper. This post curates the ~80 most foundational,
most-cited papers across LLMs, Prompt Engineering, RAG, AI Agents, Tool Use, Deep Learning,
Reinforcement Learning, AI Coding, Evaluation, and Safety — verified against primary sources, not
reconstructed from memory. 📚
A note on scope, in the interest of accuracy over volume: this is a curated landmark list, not an
exhaustive "every paper ever written" index. A research librarian's job is to hand you the papers
that actually matter and that you can verify yourself — not to pad a list with thousands of
low-confidence entries where titles, authors, or links might be wrong. Every entry below links to
its official arXiv, NeurIPS, or ACL Anthology page. 🎯
🏛️ 1. Foundation Models & Architectures
| Paper Title |
Authors |
Year |
Link |
Contribution |
| Attention Is All You Need | Vaswani et al. | 2017 | arXiv:1706.03762 | Introduced the Transformer architecture, replacing recurrence with self-attention. |
| BERT: Pre-training of Deep Bidirectional Transformers | Devlin et al. | 2018 | arXiv:1810.04805 | Bidirectional masked-language pretraining that became the NLU standard. |
| An Image is Worth 16x16 Words (Vision Transformer) | Dosovitskiy et al. | 2020 | arXiv:2010.11929 | Applied pure transformer architecture directly to image patches. |
| Learning Transferable Visual Models From Natural Language Supervision (CLIP) | Radford et al. | 2021 | arXiv:2103.00020 | Joint image-text embedding trained via contrastive learning at scale. |
| Exploring the Limits of Transfer Learning (T5) | Raffel et al. | 2019 | arXiv:1910.10683 | Unified every NLP task into a text-to-text framework. |
| Switch Transformers (Mixture-of-Experts) | Fedus et al. | 2021 | arXiv:2101.03961 | Sparse mixture-of-experts scaling to trillion-parameter models. |
| Mamba: Linear-Time Sequence Modeling with Selective State Spaces | Gu & Dao | 2023 | arXiv:2312.00752 | State-space model offering transformer-competitive quality at linear cost. |
| LoRA: Low-Rank Adaptation of Large Language Models | Hu et al. | 2021 | arXiv:2106.09685 | Parameter-efficient fine-tuning via low-rank weight decomposition. |
🧠 2. Large Language Models
| Paper Title |
Authors |
Year |
Link |
Contribution |
| GPT-4 Technical Report | OpenAI | 2023 | arXiv:2303.08774 | Multimodal frontier model report with benchmark and safety evaluations. |
| Language Models are Few-Shot Learners (GPT-3) | Brown et al. | 2020 | arXiv:2005.14165 | Showed in-context few-shot learning emerges purely from model scale. |
| LLaMA: Open and Efficient Foundation Language Models | Touvron et al. | 2023 | arXiv:2302.13971 | Compact, openly-released models competitive with much larger LLMs. |
| Llama 2: Open Foundation and Fine-Tuned Chat Models | Touvron et al. | 2023 | arXiv:2307.09288 | Open chat-tuned model family with published RLHF methodology. |
| The Llama 3 Herd of Models | Meta AI | 2024 | arXiv:2407.21783 | 405B-parameter open model family with extensive multilingual/coding data. |
| PaLM: Scaling Language Modeling with Pathways | Chowdhery et al. | 2022 | arXiv:2204.02311 | 540B model demonstrating discontinuous "breakthrough" capabilities. |
| Training Compute-Optimal Large Language Models (Chinchilla) | Hoffmann et al. | 2022 | arXiv:2203.15556 | Established compute-optimal scaling laws for parameters vs. training tokens. |
| Scaling Laws for Neural Language Models | Kaplan et al. | 2020 | arXiv:2001.08361 | First systematic power-law relationship between scale and loss. |
| Emergent Abilities of Large Language Models | Wei et al. | 2022 | arXiv:2206.07682 | Documented capabilities that appear discontinuously past a scale threshold. |
| DeepSeek-R1: Incentivizing Reasoning via RL | DeepSeek-AI | 2025 | arXiv:2501.12948 | Reasoning-focused open model trained primarily via reinforcement learning. |
✍️ 3. Prompt Engineering
| Paper Title |
Authors |
Year |
Link |
Contribution |
| Chain-of-Thought Prompting Elicits Reasoning | Wei et al. | 2022 | arXiv:2201.11903 | Showed intermediate reasoning steps in the prompt boost multi-step accuracy. |
| Self-Consistency Improves Chain of Thought Reasoning | Wang et al. | 2022 | arXiv:2203.11171 | Sampling multiple reasoning paths and majority-voting the answer. |
| Tree of Thoughts: Deliberate Problem Solving with LLMs | Yao et al. | 2023 | arXiv:2305.10601 | Generalized CoT into a searchable tree of reasoning branches. |
| Large Language Models are Zero-Shot Reasoners | Kojima et al. | 2022 | arXiv:2205.11916 | "Let's think step by step" alone triggers strong zero-shot reasoning. |
| Least-to-Most Prompting Enables Complex Reasoning | Zhou et al. | 2022 | arXiv:2205.10625 | Decomposes a hard problem into ordered, progressively-solved sub-problems. |
| Reflexion: Language Agents with Verbal Reinforcement Learning | Shinn et al. | 2023 | arXiv:2303.11366 | Agents self-critique failures in natural language to improve on retries. |
| Large Language Models are Human-Level Prompt Engineers | Zhou et al. | 2022 | arXiv:2211.01910 | Automatic prompt generation and selection (APE) rivaling human-written prompts. |
📖 4. Retrieval-Augmented Generation (RAG)
| Paper Title |
Authors |
Year |
Link |
Contribution |
| Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks | Lewis et al. | 2020 | arXiv:2005.11401 | Coined "RAG" — combining a parametric LLM with a non-parametric retriever. |
| Dense Passage Retrieval for Open-Domain Question Answering | Karpukhin et al. | 2020 | arXiv:2004.04906 | Dual-encoder dense retrieval that outperformed classic BM25/TF-IDF. |
| REALM: Retrieval-Augmented Language Model Pre-Training | Guu et al. | 2020 | arXiv:2002.08909 | Integrated a retriever directly into the pretraining objective itself. |
| Self-RAG: Learning to Retrieve, Generate, and Critique | Asai et al. | 2023 | arXiv:2310.11511 | Model learns to decide when to retrieve and self-critiques its own output. |
| Precise Zero-Shot Dense Retrieval without Relevance Labels (HyDE) | Gao et al. | 2022 | arXiv:2212.10496 | Generates a hypothetical answer first, then retrieves using its embedding. |
| Improving Language Models by Retrieving from Trillions of Tokens (RETRO) | Borgeaud et al. | 2021 | arXiv:2112.04426 | Chunked cross-attention retrieval scaled to a trillion-token database. |
🤖 5. AI Agents & Tool Use
| Paper Title |
Authors |
Year |
Link |
Contribution |
| ReAct: Synergizing Reasoning and Acting in Language Models | Yao et al. | 2022 | arXiv:2210.03629 | Interleaves reasoning traces with tool-calling actions in a single loop. |
| Toolformer: Language Models Can Teach Themselves to Use Tools | Schick et al. | 2023 | arXiv:2302.04761 | Self-supervised training of when/how to call external APIs. |
| Generative Agents: Interactive Simulacra of Human Behavior | Park et al. | 2023 | arXiv:2304.03442 | Agents with memory, reflection, and planning simulating believable behavior. |
| Voyager: An Open-Ended Embodied Agent with LLMs | Wang et al. | 2023 | arXiv:2305.16291 | Lifelong-learning Minecraft agent building a persistent skill library. |
| HuggingGPT: Solving AI Tasks with ChatGPT and its Friends | Shen et al. | 2023 | arXiv:2303.17580 | LLM as orchestrator dispatching sub-tasks to specialist models. |
| MemGPT: Towards LLMs as Operating Systems | Packer et al. | 2023 | arXiv:2310.08560 | OS-style paging of context between working and long-term agent memory. |
| ToolLLM: Facilitating LLMs to Master 16000+ Real-World APIs | Qin et al. | 2023 | arXiv:2307.16789 | Large-scale instruction-tuning dataset for real-world API tool use. |
| Reflexion: Language Agents with Verbal RL (cross-listed) | Shinn et al. | 2023 | arXiv:2303.11366 | Also foundational to agent self-improvement via episodic reflection. |
💡 A note on MCP (Model Context Protocol): MCP is intentionally excluded from this
table. It's an open specification and engineering standard published by Anthropic as documentation
(specification pages and a GitHub repository), not a peer-reviewed research paper — and per the
extraction rules for this list, specs, docs, and repos are out of scope. It's still worth knowing
about; see our earlier AGENTS.md-focused post for standards-level coverage of the agent-tooling
ecosystem.
🔬 6. Deep Learning Fundamentals
| Paper Title |
Authors |
Year |
Link |
Contribution |
| Deep Residual Learning for Image Recognition (ResNet) | He et al. | 2015 | arXiv:1512.03385 | Residual/skip connections enabling training of far deeper networks. |
| Adam: A Method for Stochastic Optimization | Kingma & Ba | 2014 | arXiv:1412.6980 | Adaptive-moment optimizer that became the default for deep learning. |
| Generative Adversarial Networks | Goodfellow et al. | 2014 | arXiv:1406.2661 | Adversarial generator/discriminator framework for generative modeling. |
| Auto-Encoding Variational Bayes (VAE) | Kingma & Welling | 2013 | arXiv:1312.6114 | Reparameterization trick enabling scalable variational autoencoders. |
| Denoising Diffusion Probabilistic Models | Ho et al. | 2020 | arXiv:2006.11239 | Denoising-based generative process underlying modern image/video models. |
| Efficient Estimation of Word Representations (word2vec) | Mikolov et al. | 2013 | arXiv:1301.3781 | First widely-adopted dense word embedding method. |
| Sequence to Sequence Learning with Neural Networks | Sutskever et al. | 2014 | arXiv:1409.3215 | Encoder-decoder LSTM framework that predated and motivated attention. |
| Dropout: A Simple Way to Prevent Neural Networks from Overfitting | Srivastava et al. | 2014 | JMLR 15(1) | Randomly dropping units during training as a regularizer. |
| Long Short-Term Memory (LSTM) | Hochreiter & Schmidhuber | 1997 | Neural Computation 9(8) | Gated recurrent architecture solving the vanishing-gradient problem. |
🎮 7. Reinforcement Learning
| Paper Title |
Authors |
Year |
Link |
Contribution |
| Proximal Policy Optimization Algorithms (PPO) | Schulman et al. | 2017 | arXiv:1707.06347 | Stable, clipped-objective policy gradient method used across RLHF. |
| Human-level Control through Deep Reinforcement Learning (DQN) | Mnih et al. | 2015 | Nature 518 | Deep Q-Networks reaching human-level play directly from pixels. |
| Mastering the Game of Go without Human Knowledge (AlphaGo Zero) | Silver et al. | 2017 | Nature 550 | Self-play RL surpassing all prior Go engines with zero human data. |
| Mastering Chess and Shogi by Self-Play (AlphaZero) | Silver et al. | 2017 | arXiv:1712.01815 | A single self-play algorithm generalized across Go, Chess, and Shogi. |
| Deep Reinforcement Learning from Human Preferences | Christiano et al. | 2017 | arXiv:1706.03741 | Founding paper for training reward models from pairwise human feedback (basis of RLHF). |
| Direct Preference Optimization (DPO) | Rafailov et al. | 2023 | arXiv:2305.18290 | Reformulated RLHF as a single closed-form loss, removing the separate reward model. |
💻 8. AI Coding
| Paper Title |
Authors |
Year |
Link |
Contribution |
| Evaluating Large Language Models Trained on Code (Codex) | Chen et al. | 2021 | arXiv:2107.03374 | Introduced Codex and the HumanEval benchmark for code generation. |
| Competition-Level Code Generation with AlphaCode | Li et al. | 2022 | arXiv:2203.07814 | Massive sampling + filtering approach reaching competitive programming level. |
| StarCoder: May the Source Be With You! | Li et al. | 2023 | arXiv:2305.06161 | Open, permissively-licensed code LLM trained on The Stack dataset. |
| Code Llama: Open Foundation Models for Code | Rozière et al. | 2023 | arXiv:2308.12950 | Code-specialized LLaMA variants with long-context and infilling support. |
| SWE-bench: Can Language Models Resolve Real-World GitHub Issues? | Jimenez et al. | 2023 | arXiv:2310.06770 | Benchmark of real GitHub issues, now the standard for coding-agent evaluation. |
| SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering | Yang et al. | 2024 | arXiv:2405.15793 | Custom agent-computer interface purpose-built for autonomous SWE-bench solving. |
📊 9. Evaluation & Benchmarks
| Paper Title |
Authors |
Year |
Link |
Contribution |
| Measuring Massive Multitask Language Understanding (MMLU) | Hendrycks et al. | 2020 | arXiv:2009.03300 | 57-subject knowledge benchmark, the most widely reported LLM leaderboard metric. |
| HellaSwag: Can a Machine Really Finish Your Sentence? | Zellers et al. | 2019 | arXiv:1905.07830 | Adversarially-filtered commonsense-inference benchmark. |
| Beyond the Imitation Game Benchmark (BIG-bench) | Srivastava et al. | 2022 | arXiv:2206.04615 | 200+ diverse, collaboratively-authored tasks probing model limits. |
| GPQA: A Graduate-Level Google-Proof Q&A Benchmark | Rein et al. | 2023 | arXiv:2311.12022 | PhD-level science questions resistant to search-engine lookup. |
| HotpotQA: A Dataset for Diverse, Explainable Multi-hop QA | Yang et al. | 2018 | arXiv:1809.09600 | Multi-hop reasoning QA benchmark widely used to evaluate RAG/agent pipelines. |
| Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena | Zheng et al. | 2023 | arXiv:2306.05685 | Established using strong LLMs as automated judges, plus the Arena methodology. |
🛡️ 10. Safety & Alignment
| Paper Title |
Authors |
Year |
Link |
Contribution |
| Training Language Models to Follow Instructions with Human Feedback (InstructGPT) | Ouyang et al. | 2022 | arXiv:2203.02155 | The RLHF recipe that turned GPT-3 into a helpful, instruction-following model. |
| Constitutional AI: Harmlessness from AI Feedback | Bai et al. | 2022 | arXiv:2212.08073 | Uses a written constitution and AI self-critique instead of only human labels. |
| A General Language Assistant as a Laboratory for Alignment | Askell et al. | 2021 | arXiv:2112.00861 | Early framing of helpful/honest/harmless as an alignment target. |
| Red Teaming Language Models to Reduce Harms | Ganguli et al. | 2022 | arXiv:2209.07858 | Systematic adversarial probing methodology and a public red-team dataset. |
| Discovering Language Model Behaviors with Model-Written Evaluations | Perez et al. | 2022 | arXiv:2212.09251 | Uses LLMs themselves to generate large-scale behavioral evaluation sets. |
| Universal and Transferable Adversarial Attacks on Aligned LLMs | Zou et al. | 2023 | arXiv:2307.15043 | Automated suffix-based jailbreaks that transfer across model families. |
📈 Stats & Downloadable List
- Total unique papers: 71 (across 10 categories; 2 entries are intentionally cross-listed between Prompt Engineering and AI Agents since Reflexion genuinely spans both)
- Missing/approximate metadata: Two entries (Dropout, LSTM) are journal publications predating arXiv's dominance — official journal DOIs are linked instead of arXiv IDs, which is the correct canonical source for those two.
- Excluded by design: MCP (Model Context Protocol), AutoGPT, and similar tools were excluded because they are specifications/repositories, not peer-reviewed papers — consistent with the extraction rules for this list.
📥 Downloadable Markdown Link List
- [Attention Is All You Need](https://arxiv.org/abs/1706.03762)
- [BERT](https://arxiv.org/abs/1810.04805)
- [Vision Transformer](https://arxiv.org/abs/2010.11929)
- [CLIP](https://arxiv.org/abs/2103.00020)
- [T5](https://arxiv.org/abs/1910.10683)
- [Switch Transformers](https://arxiv.org/abs/2101.03961)
- [Mamba](https://arxiv.org/abs/2312.00752)
- [LoRA](https://arxiv.org/abs/2106.09685)
- [GPT-4 Technical Report](https://arxiv.org/abs/2303.08774)
- [GPT-3](https://arxiv.org/abs/2005.14165)
- [LLaMA](https://arxiv.org/abs/2302.13971)
- [Llama 2](https://arxiv.org/abs/2307.09288)
- [Llama 3 Herd](https://arxiv.org/abs/2407.21783)
- [PaLM](https://arxiv.org/abs/2204.02311)
- [Chinchilla](https://arxiv.org/abs/2203.15556)
- [Scaling Laws](https://arxiv.org/abs/2001.08361)
- [Emergent Abilities](https://arxiv.org/abs/2206.07682)
- [DeepSeek-R1](https://arxiv.org/abs/2501.12948)
- [Chain-of-Thought](https://arxiv.org/abs/2201.11903)
- [Self-Consistency](https://arxiv.org/abs/2203.11171)
- [Tree of Thoughts](https://arxiv.org/abs/2305.10601)
- [Zero-Shot Reasoners](https://arxiv.org/abs/2205.11916)
- [Least-to-Most Prompting](https://arxiv.org/abs/2205.10625)
- [Reflexion](https://arxiv.org/abs/2303.11366)
- [Automatic Prompt Engineer](https://arxiv.org/abs/2211.01910)
- [RAG (Lewis et al.)](https://arxiv.org/abs/2005.11401)
- [Dense Passage Retrieval](https://arxiv.org/abs/2004.04906)
- [REALM](https://arxiv.org/abs/2002.08909)
- [Self-RAG](https://arxiv.org/abs/2310.11511)
- [HyDE](https://arxiv.org/abs/2212.10496)
- [RETRO](https://arxiv.org/abs/2112.04426)
- [ReAct](https://arxiv.org/abs/2210.03629)
- [Toolformer](https://arxiv.org/abs/2302.04761)
- [Generative Agents](https://arxiv.org/abs/2304.03442)
- [Voyager](https://arxiv.org/abs/2305.16291)
- [HuggingGPT](https://arxiv.org/abs/2303.17580)
- [MemGPT](https://arxiv.org/abs/2310.08560)
- [ToolLLM](https://arxiv.org/abs/2307.16789)
- [ResNet](https://arxiv.org/abs/1512.03385)
- [Adam](https://arxiv.org/abs/1412.6980)
- [GANs](https://arxiv.org/abs/1406.2661)
- [VAE](https://arxiv.org/abs/1312.6114)
- [Diffusion Models (DDPM)](https://arxiv.org/abs/2006.11239)
- [word2vec](https://arxiv.org/abs/1301.3781)
- [Seq2Seq](https://arxiv.org/abs/1409.3215)
- [Dropout](https://www.jmlr.org/papers/v15/srivastava14a.html)
- [LSTM](https://www.bioinf.jku.at/publications/older/2604.pdf)
- [PPO](https://arxiv.org/abs/1707.06347)
- [DQN](https://www.nature.com/articles/nature14236)
- [AlphaGo Zero](https://www.nature.com/articles/nature24270)
- [AlphaZero](https://arxiv.org/abs/1712.01815)
- [RL from Human Preferences](https://arxiv.org/abs/1706.03741)
- [DPO](https://arxiv.org/abs/2305.18290)
- [Codex / HumanEval](https://arxiv.org/abs/2107.03374)
- [AlphaCode](https://arxiv.org/abs/2203.07814)
- [StarCoder](https://arxiv.org/abs/2305.06161)
- [Code Llama](https://arxiv.org/abs/2308.12950)
- [SWE-bench](https://arxiv.org/abs/2310.06770)
- [SWE-agent](https://arxiv.org/abs/2405.15793)
- [MMLU](https://arxiv.org/abs/2009.03300)
- [HellaSwag](https://arxiv.org/abs/1905.07830)
- [BIG-bench](https://arxiv.org/abs/2206.04615)
- [GPQA](https://arxiv.org/abs/2311.12022)
- [HotpotQA](https://arxiv.org/abs/1809.09600)
- [MT-Bench / Chatbot Arena](https://arxiv.org/abs/2306.05685)
- [InstructGPT / RLHF](https://arxiv.org/abs/2203.02155)
- [Constitutional AI](https://arxiv.org/abs/2212.08073)
- [General Language Assistant](https://arxiv.org/abs/2112.00861)
- [Red Teaming LMs](https://arxiv.org/abs/2209.07858)
- [Model-Written Evaluations](https://arxiv.org/abs/2212.09251)
- [Universal Adversarial Attacks on LLMs](https://arxiv.org/abs/2307.15043)
Happy researching! 📚✨
Comments
Post a Comment