Research papers are the blueprint of modern AI: While frameworks handle execution, white papers reveal the core architectures, optimization techniques, and breakthroughs driving the next generation of LLMs and RAG systems..
📑 In This Post
1. Large Language Models
This category covers new foundation-model releases and the architectural or serving tricks that make them cheaper to run — the plumbing that everything else in this post ultimately sits on top of.
| Paper Title | Lab / Authors | Date | Link | Contribution |
|---|---|---|---|---|
| Hunyuan-A13B Technical Report | Tencent Hunyuan | Sep 23, 2026 | arXiv | An open Mixture-of-Experts model that activates only 13B of its 80B parameters per token and switches between fast and slow "thinking" modes depending on task difficulty, aiming to approach much larger models on math, code, and agent tasks at a fraction of the inference cost. |
| DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression | DeepSeek | Sep 17, 2026 | arXiv | A 552B-parameter multimodal MoE model with a new encoder-decoder split that shrinks its per-token memory footprint to roughly a quarter of its predecessor's, aimed at making long-context, agentic workloads dramatically cheaper to serve. |
| Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs | MBZUAI | Sep 22, 2026 | arXiv | A training-free speed-up for diffusion-style language models that lets the model verify its own draft tokens in parallel, easing a memory bottleneck and reportedly reaching double-digit speedups on math and code benchmarks. |
👶 Beginner's starting point
"Attention Is All You Need" (Vaswani et al., 2017) — start here because the transformer architecture it introduces is still, nine years later, the skeleton underneath essentially every model in this list.
2. Prompt Engineering
This was a genuinely quiet week for classic inference-time prompt design — the two papers that surfaced instead look at how the words in a prompt shape a model's behavior at training time and under adversarial framing.
| Paper Title | Lab / Authors | Date | Link | Contribution |
|---|---|---|---|---|
| Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training | Alibaba-PAI | Sep 14, 2026 | arXiv | Rather than treating every training prompt equally, this scores how much a prompt can still teach the model and has a teacher model rewrite the least useful ones into more informative versions, improving downstream reasoning benchmarks. |
| Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid? | Independent | Sep 13, 2026 | arXiv | A single-author study showing that a three-turn prompt escalation framed around a curious child nudges several frontier chatbots toward disclosing more speculative architectural detail than a neutral prompt would — a case study in multi-turn prompt design as a soft jailbreak vector. |
👶 Beginner's starting point
"Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (Wei et al., 2022) — the simplest prompting trick in the field (asking the model to show its work) is still the ancestor of nearly every advanced technique published since.
3. Retrieval-Augmented Generation (RAG)
Two papers this week both attack the unglamorous but expensive part of RAG: getting messy real-world content into a form a retriever can actually use well.
| Paper Title | Lab / Authors | Date | Link | Contribution |
|---|---|---|---|---|
| Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion | Yellow.ai | Sep 21, 2026 | arXiv | Extends Yellow.ai's own production chunking pipeline so any document format first renders to PDF and then converts to retrieval-ready markdown in one pass, cutting the token cost of preparing enterprise documents for RAG by roughly three-quarters. |
| Self-Evolving Search Index | Yonsei University | Sep 17, 2026 | arXiv | Lets a retrieval index rewrite the descriptive keys it uses for documents by diagnosing where retrieval is failing and testing fixes on its own, removing the need for a human to hand-tune the index as the underlying content or queries shift. |
👶 Beginner's starting point
"Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks" (Lewis et al., 2020) — the paper that coined "RAG" and laid out the retrieve-then-generate architecture that both papers above are still optimizing around.
4. AI Agents & Tool Use
By far the busiest category this window, and one clear sub-theme swallowed most of the oxygen.
| Paper Title | Lab / Authors | Date | Link | Contribution |
|---|---|---|---|---|
| SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness | NVIDIA | Sep 17, 2026 | arXiv | Scales up an auto-research loop that tests and refines coding-agent scaffolding across many diverse environments, cutting recorded token traffic by roughly 45-49% versus native Codex and Claude Code harnesses at comparable performance. |
| RRSI: Regularized Recursive Self-Improvement of Agent Harnesses | Sep 21, 2026 | arXiv | Adds explicit guardrails — a shrinking edit budget and a pruning step — to the practice of letting an agent rewrite its own harness, so the gains generalize rather than just memorizing whichever benchmark the harness was tuned against. | |
| EvoOntology: A Self-Evolving Ontology Layer for Data Agents | RUC-DataLab | Sep 14, 2026 | arXiv | Gives data agents a self-updating map of an organization's tables, files, and databases that the agent itself builds and refines, so it can look things up instead of guessing from raw column names and file paths. |
| An Empirical Study of Harness Design for Coding Agents | Zoom Communications | Sep 17, 2026 | arXiv | Systematically isolates which harness components — planning, tool choice, context management — actually help across 176 controlled settings, finding for instance that predefined tools mainly help models that are weak at using a bare shell interface. |
👶 Beginner's starting point
"ReAct: Synergizing Reasoning and Acting in Language Models" (Yao et al., 2022) — the reason-then-act loop it introduces is still the backbone pattern hiding inside most of the harnesses above.
5. Deep Learning Fundamentals
The smaller, quieter papers in this window that poke at how and why models learn what they learn, rather than chasing a leaderboard number.
| Paper Title | Lab / Authors | Date | Link | Contribution |
|---|---|---|---|---|
| Continual Learning Mechanisms Compose for Long-Horizon Memorization | Johns Hopkins | Sep 7, 2026 | arXiv | Shows that combining several complementary anti-forgetting techniques at once, rather than relying on any single one, lets a model retain far more of what it learned across 100 sequential fine-tuning tasks — lifting retention nearly 30-fold over plain sequential training. |
| The Information Geometry of Large Language Models Is Shared, Learned, and Controllable | UCL-affiliated | Sep 10, 2026 | arXiv | Argues that very different architectures — transformers, state-space models, recurrent networks — converge on a similar underlying geometry of their output predictions, and shows that geometry can be used to edit model behavior with less collateral disturbance than standard methods. |
| Disentangling Representation Evolution in Transformers through Directional Decomposition | ByteDance | Sep 14, 2026 | arXiv | Splits the updates a transformer layer makes to its hidden state into "keep the same direction" and "point somewhere new" components, and finds that suppressing the same-direction part during pretraining actually improves downstream task performance. |
👶 Beginner's starting point
"Deep Residual Learning for Image Recognition" (He et al., 2015) — the skip-connection trick that made today's very deep networks, transformers included, trainable in the first place.
6. Reinforcement Learning
The clearest sub-theme outside of agent harnesses this window: a cluster of papers all poking at the mechanics of on-policy distillation, where a smaller "student" model learns from a larger "teacher" using the student's own generated rollouts.
| Paper Title | Lab / Authors | Date | Link | Contribution |
|---|---|---|---|---|
| Bellman Policy Optimization | Apodex | Sep 14, 2026 | arXiv | A critic-free alternative to GRPO-style reinforcement learning for reasoning models that rewrites the optimization objective at the level of an entire generated trajectory instead of estimating value at every intermediate step. |
| Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening | Shanghai AI Lab | Sep 16, 2026 | arXiv | Identifies "value flattening" — where a PPO critic's estimates stay oddly flat even as true outcomes swing sharply — as an overlooked failure mode, and shows that training the critic on just a few well-chosen states per response largely fixes it. |
| When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation | Microsoft | Sep 17, 2026 | arXiv | Explains why distilled students sometimes ramble on for thousands of extra tokens: student and teacher can disagree on which specific token means "stop," and treating functionally equivalent stop tokens as one shared signal largely resolves the inflation. |
| 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation | Ant Group | Sep 21, 2026 | arXiv | Proposes a way to measure how noisy a single-sample gradient estimate is at each token during sparse on-policy distillation, and shows that selecting tokens for both usefulness and estimate reliability lets a model match full-token training while touching under 1% of tokens. |
👶 Beginner's starting point
"Proximal Policy Optimization Algorithms" (Schulman et al., 2017) — PPO is still the reference point every critic tweak and distillation trick above is implicitly compared against.
7. AI Coding
Coding-agent research this week split between building bigger, more realistic training environments and quietly worrying about whether the standard benchmarks can even be trusted anymore.
| Paper Title | Lab / Authors | Date | Link | Contribution |
|---|---|---|---|---|
| Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It? | Shanghai Jiao Tong U. | Aug 21, 2026 | arXiv | Tests whether coding agents actually understand a repository or just recognize it, by transforming SWE-bench repos at evaluation time — renaming, reordering, rewriting code while preserving behavior — and finds performance drops noticeably once familiar surface cues are erased. |
| CodeMidas: Scaling Agentic Coding RL Environments from Code Itself | Xiaomi MiMo | Sep 18, 2026 | arXiv | Builds reinforcement-learning training tasks directly out of existing open-source codebases' own functionality, rather than scraped GitHub issues, producing over 5,500 verified tasks across 23 languages that measurably improve a coding agent's issue-repair and terminal-work skills. |
| One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents | Alibaba | Sep 20, 2026 | arXiv | Trains separate expert models on different categories of software-engineering tasks to avoid the "whack-a-mole" problem where progress on one task type quietly regresses another, then merges the experts back into a single deployable model via distillation. |
👶 Beginner's starting point
"Evaluating Large Language Models Trained on Code" (Chen et al., 2021) — the Codex/HumanEval paper that established "can it actually run and pass" as the bar for judging code-generating models, a bar this week's memorization paper is directly interrogating.
8. Evaluation & Benchmarks
A smaller but pointed set of papers this window, all asking some version of "can we actually trust the number this benchmark just gave us?"
| Paper Title | Lab / Authors | Date | Link | Contribution |
|---|---|---|---|---|
| WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents | Carnegie Mellon U. | Sep 23, 2026 | arXiv | A new benchmark that checks whether AI research agents can correctly predict how changing one component of an experiment will affect the outcome, rather than just running experiments and reporting whatever number comes out. |
| Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation | UC Santa Cruz | Sep 17, 2026 | arXiv | Borrows small-area statistical estimation techniques from survey research to get much tighter per-category performance estimates from a limited evaluation budget, plus a way to check which estimator to trust without needing a separate validation sample. |
| Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles | Nat'l U. of Singapore | Sep 2, 2026 | arXiv | Uses systematic code-mutation testing to show that the correctness checkers behind popular GPU-kernel benchmarks silently miss about one in six real bugs, especially precision errors — a real problem now that those same checkers feed reinforcement-learning reward signals. |
👶 Beginner's starting point
"SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (Jimenez et al., 2023) — the benchmark that this window's coding papers are either using, stress-testing, or actively trying to patch the leakage problems of.
📝 Summary
- Large Language Models: three releases focused on doing more with less — cheaper KV caches, adaptive reasoning depth, faster diffusion-style decoding.
- Prompt Engineering: a quiet week — prompt-shaped training curricula and a cautionary multi-turn jailbreak study, rather than new inference-time techniques.
- RAG: both papers attack messy document ingestion and self-tuning indexes rather than the generation side.
- AI Agents & Tool Use: the week's dominant story — agent harnesses that rewrite themselves, and the guardrails researchers are racing to add.
- Deep Learning Fundamentals: smaller, mechanistic papers on forgetting, representation geometry, and how transformer layers actually update state.
- Reinforcement Learning: a cluster of papers debugging the fine mechanics of on-policy distillation — stop-token mismatches, noisy gradients, flat critics.
- AI Coding: bigger RL training environments built from real code, plus a pointed challenge to whether SWE-bench-style benchmarks are measuring understanding or memorization.
- Evaluation & Benchmarks: a small but sharp set of papers asking whether the graders themselves — and the statistics behind limited eval budgets — can be trusted.
Comments
Post a Comment