Research index
Every arXiv paper matching a fixed set of terms on agentic engineering, tagged so you can filter them yourself instead of trusting a ranking.
- Harness-Zero: Harness Distillation via Agent-as-Harness
2026-09-21
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
2026-09-21
- Emergent Collusion in Long-Horizon LLM Agent Interaction
2026-09-21
- ActGov: Governing LLM Agent Actions via Policy-Constrained Validation
2026-09-21
- Dissecting Agentic Forensics: The Role of Triage, Prompting, and Evidence Arbitration in Open-World Fake Image Detection
2026-09-21
- Few-Shot Demonstrations Elicit the Use of In-Context World Representations in LLMs
2026-09-21
- A Lean and Spec-Driven AI-Assisted Software Development Lifecycle for Applied AI Education: The AI-SDLC Approach
2026-09-21
- TTSE: A Two-Track Online Self-Evolution Framework
2026-09-21
- APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction
2026-09-21
- MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM Agents
2026-09-21
- Self-Healing Harness for Runtime Oversight of Agent Self-Modification
2026-09-21
- EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation
2026-09-21
- Connecting the Dots in Agentic AI Security: A Cross-Dimensional Threat Taxonomy, Evaluation Maturity, and Open Challenges
2026-09-20
- SyzHarness: Patch-Based Kernel Bug Reproduction with LLM-Synthesized Fuzzing Harnesses
2026-09-20
- FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model
2026-09-20
- BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents
2026-09-20
- RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents
2026-09-20
- PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents
2026-09-20
- Human-guided physics-constrained AI agents construct an auditable model of soil-plug evolution
2026-09-20
- CTRL: Control-Based Time Series Forecasting with LLM-Guided Residual Learning
2026-09-20
- Automatic multimodal UX improvement recommendations from LLM agent user simulations
2026-09-19
- The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI
2026-09-19
- The Price of Safety: Benign-Case Utility and Token Overhead of Memory-Poisoning Defenses in LLM Agents
2026-09-19
- SelfOp: An Optimization Algorithm for Self-Improving Security Agents
2026-09-19
- Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems
2026-09-19
- Vision2CAD: A Visual Agent Harness for Explicit Geometry Referencing and Localization in Parametric CAD Modeling
2026-09-19
- Zero-Trust Authorization and Discovery for Enterprise MCP
2026-09-18
- Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents
2026-09-18
- AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory
2026-09-18
- An Agentic Just-in-Time Adaptive Intervention System for Personalized Sleep Support: Proof-of-Concept Study with N of 1 Data
2026-09-18
- CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents
2026-09-18
- ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL
2026-09-18
- Efficient Benchmarking in Production: A Study of an Evolving LLM Agent
2026-09-18
- Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale
2026-09-18
- Two's a Crowd: Human and AI-Based Copresence for Developers with ADHD
2026-09-18
- Can Agents Design Better Chips with a Higher Level Abstraction?
2026-09-17
- Chronicle: Cut-Point Replay for Regression Testing of LLM Agents
2026-09-17
- Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape
2026-09-17
- Language-model groups overstate consensus when replaying human deliberation on a reasoning task
2026-09-17
- How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents
2026-09-17
- A Proposal for an Agentic AI Architecture to Support Multi-Domain Decision-Making in the Brazilian Armed Forces
2026-09-17
- Neuro-Symbolic Agentic AI for Networked Low-Altitude UAVs
2026-09-17
- A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents
2026-09-17
- Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization
2026-09-17
- Rethinking Multi-Agent Collaboration: When More Is Less
2026-09-17
- AutoData: Agentic Search for Pre-training Data Selection
2026-09-17
- SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes
2026-09-17
- Self-Evolving Search Index
2026-09-17
- A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems
2026-09-17
- Closed-World Resolution Against Tool Hallucination in LLM Agents
2026-09-16
Tags are assigned by a classifier, not by hand, and their accuracy has not been measured yet. Treat them as a way to narrow the list, not as fact.