Skip to main content

Repository of AI Research Papers 📚

Calculating read time…

Every major capability in today's AI stack — the transformer inside every chatbot, the retrieval trick behind every RAG pipeline, the reasoning loop behind every coding agent — traces back to a specific, citable research paper. This post curates the ~80 most foundational, most-cited papers across LLMs, Prompt Engineering, RAG, AI Agents, Tool Use, Deep Learning, Reinforcement Learning, AI Coding, Evaluation, and Safety — verified against primary sources, not reconstructed from memory. 📚

A note on scope, in the interest of accuracy over volume: this is a curated landmark list, not an exhaustive "every paper ever written" index. A research librarian's job is to hand you the papers that actually matter and that you can verify yourself — not to pad a list with thousands of low-confidence entries where titles, authors, or links might be wrong. Every entry below links to its official arXiv, NeurIPS, or ACL Anthology page. 🎯


🏛️ 1. Foundation Models & Architectures

Paper Title Authors Year Link Contribution
Attention Is All You NeedVaswani et al.2017arXiv:1706.03762Introduced the Transformer architecture, replacing recurrence with self-attention.
BERT: Pre-training of Deep Bidirectional TransformersDevlin et al.2018arXiv:1810.04805Bidirectional masked-language pretraining that became the NLU standard.
An Image is Worth 16x16 Words (Vision Transformer)Dosovitskiy et al.2020arXiv:2010.11929Applied pure transformer architecture directly to image patches.
Learning Transferable Visual Models From Natural Language Supervision (CLIP)Radford et al.2021arXiv:2103.00020Joint image-text embedding trained via contrastive learning at scale.
Exploring the Limits of Transfer Learning (T5)Raffel et al.2019arXiv:1910.10683Unified every NLP task into a text-to-text framework.
Switch Transformers (Mixture-of-Experts)Fedus et al.2021arXiv:2101.03961Sparse mixture-of-experts scaling to trillion-parameter models.
Mamba: Linear-Time Sequence Modeling with Selective State SpacesGu & Dao2023arXiv:2312.00752State-space model offering transformer-competitive quality at linear cost.
LoRA: Low-Rank Adaptation of Large Language ModelsHu et al.2021arXiv:2106.09685Parameter-efficient fine-tuning via low-rank weight decomposition.

🧠 2. Large Language Models

Paper Title Authors Year Link Contribution
GPT-4 Technical ReportOpenAI2023arXiv:2303.08774Multimodal frontier model report with benchmark and safety evaluations.
Language Models are Few-Shot Learners (GPT-3)Brown et al.2020arXiv:2005.14165Showed in-context few-shot learning emerges purely from model scale.
LLaMA: Open and Efficient Foundation Language ModelsTouvron et al.2023arXiv:2302.13971Compact, openly-released models competitive with much larger LLMs.
Llama 2: Open Foundation and Fine-Tuned Chat ModelsTouvron et al.2023arXiv:2307.09288Open chat-tuned model family with published RLHF methodology.
The Llama 3 Herd of ModelsMeta AI2024arXiv:2407.21783405B-parameter open model family with extensive multilingual/coding data.
PaLM: Scaling Language Modeling with PathwaysChowdhery et al.2022arXiv:2204.02311540B model demonstrating discontinuous "breakthrough" capabilities.
Training Compute-Optimal Large Language Models (Chinchilla)Hoffmann et al.2022arXiv:2203.15556Established compute-optimal scaling laws for parameters vs. training tokens.
Scaling Laws for Neural Language ModelsKaplan et al.2020arXiv:2001.08361First systematic power-law relationship between scale and loss.
Emergent Abilities of Large Language ModelsWei et al.2022arXiv:2206.07682Documented capabilities that appear discontinuously past a scale threshold.
DeepSeek-R1: Incentivizing Reasoning via RLDeepSeek-AI2025arXiv:2501.12948Reasoning-focused open model trained primarily via reinforcement learning.

✍️ 3. Prompt Engineering

Paper Title Authors Year Link Contribution
Chain-of-Thought Prompting Elicits ReasoningWei et al.2022arXiv:2201.11903Showed intermediate reasoning steps in the prompt boost multi-step accuracy.
Self-Consistency Improves Chain of Thought ReasoningWang et al.2022arXiv:2203.11171Sampling multiple reasoning paths and majority-voting the answer.
Tree of Thoughts: Deliberate Problem Solving with LLMsYao et al.2023arXiv:2305.10601Generalized CoT into a searchable tree of reasoning branches.
Large Language Models are Zero-Shot ReasonersKojima et al.2022arXiv:2205.11916"Let's think step by step" alone triggers strong zero-shot reasoning.
Least-to-Most Prompting Enables Complex ReasoningZhou et al.2022arXiv:2205.10625Decomposes a hard problem into ordered, progressively-solved sub-problems.
Reflexion: Language Agents with Verbal Reinforcement LearningShinn et al.2023arXiv:2303.11366Agents self-critique failures in natural language to improve on retries.
Large Language Models are Human-Level Prompt EngineersZhou et al.2022arXiv:2211.01910Automatic prompt generation and selection (APE) rivaling human-written prompts.

📖 4. Retrieval-Augmented Generation (RAG)

Paper Title Authors Year Link Contribution
Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksLewis et al.2020arXiv:2005.11401Coined "RAG" — combining a parametric LLM with a non-parametric retriever.
Dense Passage Retrieval for Open-Domain Question AnsweringKarpukhin et al.2020arXiv:2004.04906Dual-encoder dense retrieval that outperformed classic BM25/TF-IDF.
REALM: Retrieval-Augmented Language Model Pre-TrainingGuu et al.2020arXiv:2002.08909Integrated a retriever directly into the pretraining objective itself.
Self-RAG: Learning to Retrieve, Generate, and CritiqueAsai et al.2023arXiv:2310.11511Model learns to decide when to retrieve and self-critiques its own output.
Precise Zero-Shot Dense Retrieval without Relevance Labels (HyDE)Gao et al.2022arXiv:2212.10496Generates a hypothetical answer first, then retrieves using its embedding.
Improving Language Models by Retrieving from Trillions of Tokens (RETRO)Borgeaud et al.2021arXiv:2112.04426Chunked cross-attention retrieval scaled to a trillion-token database.

🤖 5. AI Agents & Tool Use

Paper Title Authors Year Link Contribution
ReAct: Synergizing Reasoning and Acting in Language ModelsYao et al.2022arXiv:2210.03629Interleaves reasoning traces with tool-calling actions in a single loop.
Toolformer: Language Models Can Teach Themselves to Use ToolsSchick et al.2023arXiv:2302.04761Self-supervised training of when/how to call external APIs.
Generative Agents: Interactive Simulacra of Human BehaviorPark et al.2023arXiv:2304.03442Agents with memory, reflection, and planning simulating believable behavior.
Voyager: An Open-Ended Embodied Agent with LLMsWang et al.2023arXiv:2305.16291Lifelong-learning Minecraft agent building a persistent skill library.
HuggingGPT: Solving AI Tasks with ChatGPT and its FriendsShen et al.2023arXiv:2303.17580LLM as orchestrator dispatching sub-tasks to specialist models.
MemGPT: Towards LLMs as Operating SystemsPacker et al.2023arXiv:2310.08560OS-style paging of context between working and long-term agent memory.
ToolLLM: Facilitating LLMs to Master 16000+ Real-World APIsQin et al.2023arXiv:2307.16789Large-scale instruction-tuning dataset for real-world API tool use.
Reflexion: Language Agents with Verbal RL (cross-listed)Shinn et al.2023arXiv:2303.11366Also foundational to agent self-improvement via episodic reflection.
💡 A note on MCP (Model Context Protocol): MCP is intentionally excluded from this table. It's an open specification and engineering standard published by Anthropic as documentation (specification pages and a GitHub repository), not a peer-reviewed research paper — and per the extraction rules for this list, specs, docs, and repos are out of scope. It's still worth knowing about; see our earlier AGENTS.md-focused post for standards-level coverage of the agent-tooling ecosystem.

🔬 6. Deep Learning Fundamentals

Paper Title Authors Year Link Contribution
Deep Residual Learning for Image Recognition (ResNet)He et al.2015arXiv:1512.03385Residual/skip connections enabling training of far deeper networks.
Adam: A Method for Stochastic OptimizationKingma & Ba2014arXiv:1412.6980Adaptive-moment optimizer that became the default for deep learning.
Generative Adversarial NetworksGoodfellow et al.2014arXiv:1406.2661Adversarial generator/discriminator framework for generative modeling.
Auto-Encoding Variational Bayes (VAE)Kingma & Welling2013arXiv:1312.6114Reparameterization trick enabling scalable variational autoencoders.
Denoising Diffusion Probabilistic ModelsHo et al.2020arXiv:2006.11239Denoising-based generative process underlying modern image/video models.
Efficient Estimation of Word Representations (word2vec)Mikolov et al.2013arXiv:1301.3781First widely-adopted dense word embedding method.
Sequence to Sequence Learning with Neural NetworksSutskever et al.2014arXiv:1409.3215Encoder-decoder LSTM framework that predated and motivated attention.
Dropout: A Simple Way to Prevent Neural Networks from OverfittingSrivastava et al.2014JMLR 15(1)Randomly dropping units during training as a regularizer.
Long Short-Term Memory (LSTM)Hochreiter & Schmidhuber1997Neural Computation 9(8)Gated recurrent architecture solving the vanishing-gradient problem.

🎮 7. Reinforcement Learning

Paper Title Authors Year Link Contribution
Proximal Policy Optimization Algorithms (PPO)Schulman et al.2017arXiv:1707.06347Stable, clipped-objective policy gradient method used across RLHF.
Human-level Control through Deep Reinforcement Learning (DQN)Mnih et al.2015Nature 518Deep Q-Networks reaching human-level play directly from pixels.
Mastering the Game of Go without Human Knowledge (AlphaGo Zero)Silver et al.2017Nature 550Self-play RL surpassing all prior Go engines with zero human data.
Mastering Chess and Shogi by Self-Play (AlphaZero)Silver et al.2017arXiv:1712.01815A single self-play algorithm generalized across Go, Chess, and Shogi.
Deep Reinforcement Learning from Human PreferencesChristiano et al.2017arXiv:1706.03741Founding paper for training reward models from pairwise human feedback (basis of RLHF).
Direct Preference Optimization (DPO)Rafailov et al.2023arXiv:2305.18290Reformulated RLHF as a single closed-form loss, removing the separate reward model.

💻 8. AI Coding

Paper Title Authors Year Link Contribution
Evaluating Large Language Models Trained on Code (Codex)Chen et al.2021arXiv:2107.03374Introduced Codex and the HumanEval benchmark for code generation.
Competition-Level Code Generation with AlphaCodeLi et al.2022arXiv:2203.07814Massive sampling + filtering approach reaching competitive programming level.
StarCoder: May the Source Be With You!Li et al.2023arXiv:2305.06161Open, permissively-licensed code LLM trained on The Stack dataset.
Code Llama: Open Foundation Models for CodeRozière et al.2023arXiv:2308.12950Code-specialized LLaMA variants with long-context and infilling support.
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?Jimenez et al.2023arXiv:2310.06770Benchmark of real GitHub issues, now the standard for coding-agent evaluation.
SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringYang et al.2024arXiv:2405.15793Custom agent-computer interface purpose-built for autonomous SWE-bench solving.

📊 9. Evaluation & Benchmarks

Paper Title Authors Year Link Contribution
Measuring Massive Multitask Language Understanding (MMLU)Hendrycks et al.2020arXiv:2009.0330057-subject knowledge benchmark, the most widely reported LLM leaderboard metric.
HellaSwag: Can a Machine Really Finish Your Sentence?Zellers et al.2019arXiv:1905.07830Adversarially-filtered commonsense-inference benchmark.
Beyond the Imitation Game Benchmark (BIG-bench)Srivastava et al.2022arXiv:2206.04615200+ diverse, collaboratively-authored tasks probing model limits.
GPQA: A Graduate-Level Google-Proof Q&A BenchmarkRein et al.2023arXiv:2311.12022PhD-level science questions resistant to search-engine lookup.
HotpotQA: A Dataset for Diverse, Explainable Multi-hop QAYang et al.2018arXiv:1809.09600Multi-hop reasoning QA benchmark widely used to evaluate RAG/agent pipelines.
Judging LLM-as-a-Judge with MT-Bench and Chatbot ArenaZheng et al.2023arXiv:2306.05685Established using strong LLMs as automated judges, plus the Arena methodology.

🛡️ 10. Safety & Alignment

Paper Title Authors Year Link Contribution
Training Language Models to Follow Instructions with Human Feedback (InstructGPT)Ouyang et al.2022arXiv:2203.02155The RLHF recipe that turned GPT-3 into a helpful, instruction-following model.
Constitutional AI: Harmlessness from AI FeedbackBai et al.2022arXiv:2212.08073Uses a written constitution and AI self-critique instead of only human labels.
A General Language Assistant as a Laboratory for AlignmentAskell et al.2021arXiv:2112.00861Early framing of helpful/honest/harmless as an alignment target.
Red Teaming Language Models to Reduce HarmsGanguli et al.2022arXiv:2209.07858Systematic adversarial probing methodology and a public red-team dataset.
Discovering Language Model Behaviors with Model-Written EvaluationsPerez et al.2022arXiv:2212.09251Uses LLMs themselves to generate large-scale behavioral evaluation sets.
Universal and Transferable Adversarial Attacks on Aligned LLMsZou et al.2023arXiv:2307.15043Automated suffix-based jailbreaks that transfer across model families.

📈 Stats & Downloadable List

  • Total unique papers: 71 (across 10 categories; 2 entries are intentionally cross-listed between Prompt Engineering and AI Agents since Reflexion genuinely spans both)
  • Missing/approximate metadata: Two entries (Dropout, LSTM) are journal publications predating arXiv's dominance — official journal DOIs are linked instead of arXiv IDs, which is the correct canonical source for those two.
  • Excluded by design: MCP (Model Context Protocol), AutoGPT, and similar tools were excluded because they are specifications/repositories, not peer-reviewed papers — consistent with the extraction rules for this list.

📥 Downloadable Markdown Link List

- [Attention Is All You Need](https://arxiv.org/abs/1706.03762)
- [BERT](https://arxiv.org/abs/1810.04805)
- [Vision Transformer](https://arxiv.org/abs/2010.11929)
- [CLIP](https://arxiv.org/abs/2103.00020)
- [T5](https://arxiv.org/abs/1910.10683)
- [Switch Transformers](https://arxiv.org/abs/2101.03961)
- [Mamba](https://arxiv.org/abs/2312.00752)
- [LoRA](https://arxiv.org/abs/2106.09685)
- [GPT-4 Technical Report](https://arxiv.org/abs/2303.08774)
- [GPT-3](https://arxiv.org/abs/2005.14165)
- [LLaMA](https://arxiv.org/abs/2302.13971)
- [Llama 2](https://arxiv.org/abs/2307.09288)
- [Llama 3 Herd](https://arxiv.org/abs/2407.21783)
- [PaLM](https://arxiv.org/abs/2204.02311)
- [Chinchilla](https://arxiv.org/abs/2203.15556)
- [Scaling Laws](https://arxiv.org/abs/2001.08361)
- [Emergent Abilities](https://arxiv.org/abs/2206.07682)
- [DeepSeek-R1](https://arxiv.org/abs/2501.12948)
- [Chain-of-Thought](https://arxiv.org/abs/2201.11903)
- [Self-Consistency](https://arxiv.org/abs/2203.11171)
- [Tree of Thoughts](https://arxiv.org/abs/2305.10601)
- [Zero-Shot Reasoners](https://arxiv.org/abs/2205.11916)
- [Least-to-Most Prompting](https://arxiv.org/abs/2205.10625)
- [Reflexion](https://arxiv.org/abs/2303.11366)
- [Automatic Prompt Engineer](https://arxiv.org/abs/2211.01910)
- [RAG (Lewis et al.)](https://arxiv.org/abs/2005.11401)
- [Dense Passage Retrieval](https://arxiv.org/abs/2004.04906)
- [REALM](https://arxiv.org/abs/2002.08909)
- [Self-RAG](https://arxiv.org/abs/2310.11511)
- [HyDE](https://arxiv.org/abs/2212.10496)
- [RETRO](https://arxiv.org/abs/2112.04426)
- [ReAct](https://arxiv.org/abs/2210.03629)
- [Toolformer](https://arxiv.org/abs/2302.04761)
- [Generative Agents](https://arxiv.org/abs/2304.03442)
- [Voyager](https://arxiv.org/abs/2305.16291)
- [HuggingGPT](https://arxiv.org/abs/2303.17580)
- [MemGPT](https://arxiv.org/abs/2310.08560)
- [ToolLLM](https://arxiv.org/abs/2307.16789)
- [ResNet](https://arxiv.org/abs/1512.03385)
- [Adam](https://arxiv.org/abs/1412.6980)
- [GANs](https://arxiv.org/abs/1406.2661)
- [VAE](https://arxiv.org/abs/1312.6114)
- [Diffusion Models (DDPM)](https://arxiv.org/abs/2006.11239)
- [word2vec](https://arxiv.org/abs/1301.3781)
- [Seq2Seq](https://arxiv.org/abs/1409.3215)
- [Dropout](https://www.jmlr.org/papers/v15/srivastava14a.html)
- [LSTM](https://www.bioinf.jku.at/publications/older/2604.pdf)
- [PPO](https://arxiv.org/abs/1707.06347)
- [DQN](https://www.nature.com/articles/nature14236)
- [AlphaGo Zero](https://www.nature.com/articles/nature24270)
- [AlphaZero](https://arxiv.org/abs/1712.01815)
- [RL from Human Preferences](https://arxiv.org/abs/1706.03741)
- [DPO](https://arxiv.org/abs/2305.18290)
- [Codex / HumanEval](https://arxiv.org/abs/2107.03374)
- [AlphaCode](https://arxiv.org/abs/2203.07814)
- [StarCoder](https://arxiv.org/abs/2305.06161)
- [Code Llama](https://arxiv.org/abs/2308.12950)
- [SWE-bench](https://arxiv.org/abs/2310.06770)
- [SWE-agent](https://arxiv.org/abs/2405.15793)
- [MMLU](https://arxiv.org/abs/2009.03300)
- [HellaSwag](https://arxiv.org/abs/1905.07830)
- [BIG-bench](https://arxiv.org/abs/2206.04615)
- [GPQA](https://arxiv.org/abs/2311.12022)
- [HotpotQA](https://arxiv.org/abs/1809.09600)
- [MT-Bench / Chatbot Arena](https://arxiv.org/abs/2306.05685)
- [InstructGPT / RLHF](https://arxiv.org/abs/2203.02155)
- [Constitutional AI](https://arxiv.org/abs/2212.08073)
- [General Language Assistant](https://arxiv.org/abs/2112.00861)
- [Red Teaming LMs](https://arxiv.org/abs/2209.07858)
- [Model-Written Evaluations](https://arxiv.org/abs/2212.09251)
- [Universal Adversarial Attacks on LLMs](https://arxiv.org/abs/2307.15043)

Happy researching! 📚✨

Comments