Skip to main content

AI Research Papers published in July and August 2026

Calculating read time…

July and August 2026 were two of the busiest months on record for frontier AI research — a 2.8-trillion-parameter open model, a fresh Gemma generation, and a wave of papers rethinking agent harnesses, retrieval, and reasoning transparency all landed within weeks of each other. This post curates the papers from those two months that are getting the most attention right now — measured by Hugging Face community upvotes, arXiv listing activity, and citation by trackers like Papers with Code — across the same ten categories as our landmark-papers post. 📚

A note on how "famous" works for papers that are only weeks old: none of these have had time to accumulate the thousands of citations that made "Attention Is All You Need" or "Chain-of-Thought Prompting" landmarks. What we curated instead is current attention — papers the research community is actively upvoting, discussing, and building on right now. Every entry links to its official arXiv page, and every entry was verified against that page directly (not reconstructed from memory). Because recency and lasting importance are different things, each section below closes with a 👶 beginner's starting point — the older, foundational paper in that category worth reading first, before diving into this month's frontier. 🎯


🏛️ 1. Foundation Models & Architectures

Paper Title Lab / Authors Date Link Contribution
Kimi K3: Open Frontier IntelligenceMoonshot AIJul 27, 2026arXiv:2607.24653The largest open-weight model to date: a 2.8T-parameter MoE with Kimi Delta Attention and Stable LatentMoE, ~2.5x more scaling-efficient than its predecessor.
Gemma 4 Technical ReportGoogle DeepMindJul 2026arXiv:2607.02770Latest generation of Google's open-weight model family, detailing architecture and training updates over Gemma 3.
MOSS-VL Technical ReportOpenMOSSAug 2026arXiv:2608.15045Open vision-language foundation model release from the OpenMOSS team.
Motif 3: Technical ReportMotif TechnologiesAug 2026arXiv:2608.09119Third-generation foundation model report from Motif, covering architecture and training recipe changes.
xHC: Expanded Hyper-Connectionsdots studioJul 2026arXiv:2607.14530Extends hyper-connections (a residual-stream alternative) with a wider, more expressive connectivity pattern.
BDH-CQ: In-Context Learning with Recurrent Latent ReasoningPathwayAug 2026arXiv:2608.09888Explores a recurrent, latent-reasoning alternative to standard transformer in-context learning.
👶 Beginner's starting point: before these, read "Attention Is All You Need" (Vaswani et al., 2017) — the Transformer paper every architecture above still builds on.

🧠 2. Large Language Models

Paper Title Lab / Authors Date Link Contribution
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD65 authors (SLAI)Jul 2026arXiv:2607.20145Documents full-parameter post-training of DeepSeek-V4-scale models on domestic Ascend accelerators.
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRAMind LabAug 2026arXiv:2608.09819Open continual-learning LLM that self-improves post-deployment via a mixture-of-LoRA mechanism.
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning47 authors (Intern)Aug 2026arXiv:2608.14290Separates a model's factual-knowledge store from its reasoning machinery as distinct trainable components.
DiffusionGemma Technical ReportGoogle DeepMindAug 2026arXiv:2608.00146A diffusion-based language model variant of the Gemma family, an alternative to autoregressive decoding.
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory7 authorsJul 2026arXiv:2607.27919Proposes a pretrained decoder module that gives LLMs a parametric (non-retrieval) long-term memory.
👶 Beginner's starting point: "Language Models are Few-Shot Learners" (GPT-3, Brown et al., 2020) — where the modern LLM story really starts.

✍️ 3. Prompt Engineering

Paper Title Lab / Authors Date Link Contribution
Not All LLM Reasoning is Visible in the Chain-of-ThoughtIndependent / multi-lab evalJul 2026arXiv:2607.22925Shows filler tokens let frontier models improve performance via reasoning that never appears in the visible chain-of-thought.
Agent Against Agent: An Agentic System for Automatic Prompt Injection Red TeamingMultiple authorsAug 2026arXiv:2608.05108Uses one LLM agent to automatically craft and test prompt-injection attacks against another, for red-teaming at scale.
AISPA: User-Centric System Prompt Auditing for Large Language Model ApplicationsStanford UniversityJul 2026arXiv:2607.28617Framework letting end users audit what's actually in an application's hidden system prompt.
Demystifying Agent Skills: Why They Work-Until They Don'tUC San DiegoAug 2026arXiv:2608.14036Controlled study of when packaged "skill" prompts help agents at inference time, and where they silently fail.
👶 Beginner's starting point: "Chain-of-Thought Prompting Elicits Reasoning" (Wei et al., 2022) — the paper this month's CoT-transparency work is directly responding to.

📖 4. Retrieval-Augmented Generation (RAG)

Paper Title Lab / Authors Date Link Contribution
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigmsmuset.aiJul 2026arXiv:2607.26497Large-scale comparison finding classic sparse retrieval (BM25) remains highly competitive with dense/RAG pipelines as corpora scale.
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLMNovosibirsk State UniversityJul 2026arXiv:2607.11683Multi-hop, graph-structured RAG engine paired with a small domain-tuned LLM for efficient deployment.
A New Role for Relevance: Guiding Corpus Interaction in Agentic SearchTencentJul 2026arXiv:2607.24223Rethinks how relevance signals should steer an agent's turn-by-turn interaction with a retrieval corpus, not just final ranking.
Source-Aware Reranking for RAG: A Reliability Prior ApproachMilwaukee School of EngineeringJul 2026arXiv:2607.22584Reweights retrieval scores by source credibility, improving precision and adversarial robustness in a health-domain testbed.
👶 Beginner's starting point: "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks" (Lewis et al., 2020) — the paper that coined "RAG."

🤖 5. AI Agents & Tool Use

Paper Title Lab / Authors Date Link Contribution
Qwen-UI-Agent Technical ReportTongyi Lab (Alibaba)Jul 2026arXiv:2607.28227Next-generation, real-world-centric foundation GUI agent for operating computer and phone interfaces.
Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and EditableTencent HunyuanJul 2026arXiv:2607.13285Proposes structuring the ever-growing "harness" code around agents so it stays maintainable as it evolves.
StateM: Reaching 95.3% Raw Accuracy on Terminal-Bench 2.1 via Harness Scaling4 authorsAug 2026arXiv:2608.15089Shows scaling the agent harness itself — not just the model — drives large jumps on long-horizon terminal tasks.
DarwinX: Evolving Agent Harnesses Through Natural SelectionSalesforce AI ResearchAug 2026arXiv:2608.07545Applies an evolutionary search process to automatically discover better agent harness designs.
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic DesignMeituanAug 2026arXiv:2608.13560A "meta" layer that optimizes the agent harness itself for open-ended, long-horizon design tasks.
💡 What's an "agent harness"? It's the scaffolding around a model — the tool definitions, memory, retry logic, and control loop — that turns a raw LLM into a working agent. It's the single most active sub-theme in agent research this summer, hence how many of the above papers focus on scaling or evolving it rather than the underlying model.
👶 Beginner's starting point: "ReAct: Synergizing Reasoning and Acting in Language Models" (Yao et al., 2022) — the original reason-then-act agent loop.

🔬 6. Deep Learning Fundamentals

Paper Title Lab / Authors Date Link Contribution
BDH-CQ: In-Context Learning with Recurrent Latent ReasoningPathwayAug 2026arXiv:2608.09888A recurrent, brain-inspired architecture explored as a fundamental alternative to attention-based in-context learning.
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern OptimizersOpenDataLabJul 2026arXiv:2607.04033Systematic taxonomy and benchmark of the post-Adam generation of optimizers (Muon, Lion, and others) reshaping training.
Hierarchical Sparse Attention Done Right: Toward Infinite Context ModelingTencent HunyuanJul 2026arXiv:2607.02980Redesigns hierarchical sparse attention to remove the quality loss typically traded for near-infinite context length.
xHC: Expanded Hyper-Connectionsdots studioJul 2026arXiv:2607.14530Widens the hyper-connections residual mechanism, a fundamental building-block change applicable beyond any one model.
👶 Beginner's starting point: "Adam: A Method for Stochastic Optimization" (Kingma & Ba, 2014) — the optimizer these new papers are trying to move past.

🎮 7. Reinforcement Learning

Paper Title Lab / Authors Date Link Contribution
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards11 authorsJul 2026arXiv:2607.23802Transforms open-ended tasks into a form with built-in, self-verifiable rewards, extending RL with verifiable rewards (RLVR).
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent ReasoningRUC-AIBOXJul 2026arXiv:2607.12395Scales "zero" RL (RL from a base model, no SFT) to trillion-parameter scale and studies emergent reasoning.
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning11 authorsJul 2026arXiv:2607.14777Combines on-policy distillation with self-evolving curricula to make agentic RL training more sample-efficient.
Weak-to-Strong Generalization via Direct On-Policy DistillationBytedTsinghua-SIAJul 2026arXiv:2607.05394Studies whether a weaker teacher policy can still lift a stronger student via direct on-policy distillation.
👶 Beginner's starting point: "Proximal Policy Optimization Algorithms" (Schulman et al., 2017) — still the workhorse most of these papers build variations on.

💻 8. AI Coding

Paper Title Lab / Authors Date Link Contribution
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code RefactoringByteDanceAug 2026arXiv:2608.09802Extends the SWE-bench family to large-scale, multilingual refactoring tasks rather than single-issue fixes.
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution6 authorsAug 2026arXiv:2608.08311A coding agent that iteratively rewrites and improves its own core logic, with a review gate on each change.
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?OpenMOSSAug 2026arXiv:2608.19799Adapts the SWE-bench methodology to scientific-software engineering tasks rather than general GitHub issues.
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation ModelsTIGER-LabJul 2026arXiv:2607.12463A function-boundary-aware fill-in-the-middle objective used as a mid-training stage for coding-agent base models.
👶 Beginner's starting point: "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (Jimenez et al., 2023) — the benchmark every paper above is extending.

📊 9. Evaluation & Benchmarks

Paper Title Lab / Authors Date Link Contribution
ASI-Bench: At the Dawn of Artificial Superintelligence42 authorsAug 2026arXiv:2608.17271A frontier benchmark probing capabilities positioned beyond current general-intelligence tests.
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal TasksTencent HunyuanJul 2026arXiv:2607.08964Dense, reward-based grading for long-horizon terminal/command-line agent tasks, beyond pass/fail scoring.
HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationScale AIAug 2026arXiv:2608.06301Benchmarks how well LLMs can improve their own (or another model's) agent harness design.
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and DevelopmentLongCat (Meituan)Aug 2026arXiv:2608.13417Argues for evaluating AI-R&D agents on process quality over the entire run, not just the final benchmark score.
👶 Beginner's starting point: "Measuring Massive Multitask Language Understanding" (Hendrycks et al., 2020, MMLU) — still the most-cited reference point these newer, harder benchmarks are positioned against.

🛡️ 10. Safety & Alignment

Paper Title Lab / Authors Date Link Contribution
Not All LLM Reasoning is Visible in the Chain-of-ThoughtMulti-lab evaluation (13 frontier models tested)Jul 2026arXiv:2607.22925Demonstrates a concrete case of hidden reasoning that CoT-monitoring-based safety approaches would miss.
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution9 authorsAug 2026arXiv:2608.00677Automatically evolves open-ended adversarial environments to red-team agents at scale.
Stealing Reasoning Traces from Proprietary LLM APIs8 authorsAug 2026arXiv:2608.09867Shows reasoning traces can be reconstructed from black-box API access, a distillation/IP-leakage risk.
Agent Against Agent: An Agentic System for Automatic Prompt Injection Red TeamingMultiple authorsAug 2026arXiv:2608.05108Agentic system that automates prompt-injection attack generation for evaluating and hardening LLM agents.
👶 Beginner's starting point: "Training Language Models to Follow Instructions with Human Feedback" (Ouyang et al., 2022, InstructGPT) — the RLHF recipe underlying most of today's alignment work.

📈 Stats & Downloadable List

  • Total unique papers: 39 across 10 categories, all first submitted to arXiv in July or August 2026 (2 papers — "Not All LLM Reasoning is Visible in the Chain-of-Thought" and "Agent Against Agent" — are intentionally cross-listed, since they genuinely span Prompt Engineering and Safety).
  • How these were selected: pulled from arXiv's cs.CL/cs.AI/cs.LG listings for Jul–Aug 2026 and cross-checked against Hugging Face's Daily/Monthly Papers trending lists (community upvotes) and Papers with Code, then verified individually against each paper's live arXiv abstract page.
  • Excluded by design: model announcements or blog posts without an arXiv/technical-report companion (e.g. several July 2026 model launches covered only in press coverage), and anything we could not verify directly against a live arXiv page.
  • A caveat worth repeating: "trending in August 2026" and "foundational" are different claims. These papers are the ones getting attention right now — some will matter in five years, most won't. That's why every section links back to the older paper worth learning first.

📥 Downloadable Markdown Link List

- [Kimi K3: Open Frontier Intelligence](https://arxiv.org/abs/2607.24653)
- [Gemma 4 Technical Report](https://arxiv.org/abs/2607.02770)
- [MOSS-VL Technical Report](https://arxiv.org/abs/2608.15045)
- [Motif 3: Technical Report](https://arxiv.org/abs/2608.09119)
- [xHC: Expanded Hyper-Connections](https://arxiv.org/abs/2607.14530)
- [BDH-CQ: In-Context Learning with Recurrent Latent Reasoning](https://arxiv.org/abs/2608.09888)
- [SLAI T-Rex: DeepSeek-V4 Post-training on Ascend SuperPOD](https://arxiv.org/abs/2607.20145)
- [Macaron-V1: Open Continual Learning with Mixture-of-LoRA](https://arxiv.org/abs/2608.09819)
- [Intern-S2-Mobius: Decoupled Knowledge and Reasoning](https://arxiv.org/abs/2608.14290)
- [DiffusionGemma Technical Report](https://arxiv.org/abs/2608.00146)
- [Memory Decoder at Scale](https://arxiv.org/abs/2607.27919)
- [Not All LLM Reasoning is Visible in the Chain-of-Thought](https://arxiv.org/abs/2607.22925)
- [Agent Against Agent: Automatic Prompt Injection Red Teaming](https://arxiv.org/abs/2608.05108)
- [AISPA: User-Centric System Prompt Auditing](https://arxiv.org/abs/2607.28617)
- [Demystifying Agent Skills: Why They Work-Until They Don't](https://arxiv.org/abs/2608.14036)
- [BM25 Wins at Scale](https://arxiv.org/abs/2607.26497)
- [RAGU: A Multi-Step GraphRAG Engine](https://arxiv.org/abs/2607.11683)
- [A New Role for Relevance in Agentic Search](https://arxiv.org/abs/2607.24223)
- [Source-Aware Reranking for RAG](https://arxiv.org/abs/2607.22584)
- [Qwen-UI-Agent Technical Report](https://arxiv.org/abs/2607.28227)
- [Harness Handbook](https://arxiv.org/abs/2607.13285)
- [StateM: Terminal-Bench 2.1 via Harness Scaling](https://arxiv.org/abs/2608.15089)
- [DarwinX: Evolving Agent Harnesses Through Natural Selection](https://arxiv.org/abs/2608.07545)
- [AutoDesign: Meta-Harness Optimization](https://arxiv.org/abs/2608.13560)
- [OmniOpt: Taxonomy of Modern Optimizers](https://arxiv.org/abs/2607.04033)
- [Hierarchical Sparse Attention Done Right](https://arxiv.org/abs/2607.02980)
- [From RLVR to RLSVR](https://arxiv.org/abs/2607.23802)
- [Ring-Zero: Scaling Zero RL to a Trillion Parameters](https://arxiv.org/abs/2607.12395)
- [SEED: Self-Evolving On-Policy Distillation](https://arxiv.org/abs/2607.14777)
- [Weak-to-Strong Generalization via Direct On-Policy Distillation](https://arxiv.org/abs/2607.05394)
- [SWE-Bench ProMax](https://arxiv.org/abs/2608.09802)
- [Ouroboros: A Self-Developing Frontier Coding Agent](https://arxiv.org/abs/2608.08311)
- [SWE-bench Science](https://arxiv.org/abs/2608.19799)
- [Function-Aware Fill-in-the-Middle](https://arxiv.org/abs/2607.12463)
- [ASI-Bench: At the Dawn of Artificial Superintelligence](https://arxiv.org/abs/2608.17271)
- [Long-Horizon-Terminal-Bench](https://arxiv.org/abs/2607.08964)
- [HarnessOpt-Bench](https://arxiv.org/abs/2608.06301)
- [Beyond Final Scores: Evaluating Long-Horizon AI R&D Agents](https://arxiv.org/abs/2608.13417)
- [OpenART: Scaling Agent Red Teaming](https://arxiv.org/abs/2608.00677)
- [Stealing Reasoning Traces from Proprietary LLM APIs](https://arxiv.org/abs/2608.09867)

📚 References

Happy researching! 📚✨

Comments