Index de recherche

Tous les articles arXiv correspondant à une liste fixe de termes sur l'ingénierie agentique, étiquetés pour que vous puissiez les filtrer plutôt que de vous fier à un classement.

4 683 articles mis à jour 2026-09-23

  1. Harness-Zero: Harness Distillation via Agent-as-Harness

    Haoran Ye, Yuxing Lu, Haonan Dong +2 2026-09-21

    method

  2. RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

    Peng Xia, Rujun Han, Zifeng Wang +11 2026-09-21

    method code releasedharness

  3. Emergent Collusion in Long-Horizon LLM Agent Interaction

    Xinrui Shi, Yanzhe Zhang, Diyi Yang 2026-09-21

    study securitymulti-agent

  4. ActGov: Governing LLM Agent Actions via Policy-Constrained Validation

    Kaiyuan Zhang, Yuke Peng, Ke Jiang +1 2026-09-21

    method securitytool use

  5. Dissecting Agentic Forensics: The Role of Triage, Prompting, and Evidence Arbitration in Open-World Fake Image Detection

    Xianlong Li, Pietro Bongini, Niccoló Pancino +3 2026-09-21

    study incertain

  6. Few-Shot Demonstrations Elicit the Use of In-Context World Representations in LLMs

    Kohsei Matsutani, Gouki Minegishi, Core Francisco Park +3 2026-09-21

    study incertain

  7. A Lean and Spec-Driven AI-Assisted Software Development Lifecycle for Applied AI Education: The AI-SDLC Approach

    Andreas Martin, Sandro Schwander 2026-09-21

    method

  8. TTSE: A Two-Track Online Self-Evolution Framework

    Ruimin Pei, Yongkang Wu, Shangyi Zheng +5 2026-09-21

    method

  9. APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction

    Pawan K. Tripathi, Hemant Sharma, Andrew Chuang +1 2026-09-21

    method securitycode releasedharnesstool use

  10. MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM Agents

    Demetris Paschalides, Moysis Symeonides, George Pallis +1 2026-09-21

    benchmark evaluationcode releasedharnesstool use

  11. Self-Healing Harness for Runtime Oversight of Agent Self-Modification

    Sina Tayebati, Divake Kumar, Nastaran Darabi +2 2026-09-21

    method securityharness

  12. EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation

    Harshavardhan Abichandani, Penny Chong, Jiyuan Shen +7 2026-09-21

    method

  13. Connecting the Dots in Agentic AI Security: A Cross-Dimensional Threat Taxonomy, Evaluation Maturity, and Open Challenges

    Heewon Baek, Alsharif Abuadbba, Kristen Moore +2 2026-09-20

    survey securityevaluation

  14. SyzHarness: Patch-Based Kernel Bug Reproduction with LLM-Synthesized Fuzzing Harnesses

    Xingyu Li, Juefei Pu, Haonan Li +4 2026-09-20

    method harness

  15. FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model

    Jingxuan Xu, Gang Wu, Yanan Wu +13 2026-09-20

    method

  16. BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents

    Peng Kuang, Yuchun Fan, Jiangnan Li +7 2026-09-20

    benchmark evaluation

  17. RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents

    Fanyu Zhao, Ruike Cao, Liang Dong +6 2026-09-20

    method memorycode released incertain

  18. PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents

    Pirzada Suhail, Menglin Xia, Xuchao Zhang +3 2026-09-20

    method memory incertain

  19. Human-guided physics-constrained AI agents construct an auditable model of soil-plug evolution

    Jie Shi, Yimin Lu, Zhongkun Ouyang 2026-09-20

    method multi-agent

  20. CTRL: Control-Based Time Series Forecasting with LLM-Guided Residual Learning

    Minkyoung Kim, Daeun Ji, Yohan Lee +2 2026-09-20

    method incertain

  21. Automatic multimodal UX improvement recommendations from LLM agent user simulations

    Anu Chowdhury, Bin Wu, Hossein A. Rahmani +1 2026-09-19

    method incertain

  22. The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI

    Oren Perez 2026-09-19

    securitymulti-agent

  23. The Price of Safety: Benign-Case Utility and Token Overhead of Memory-Poisoning Defenses in LLM Agents

    Pritom Bhowmik 2026-09-19

    study securitymemoryevaluation

  24. SelfOp: An Optimization Algorithm for Self-Improving Security Agents

    Saad Ullah, Yigitcan Kaya, Christopher Kruegel +2 2026-09-19

    method

  25. Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems

    Fayeq Jeelani Syed, Rehan Ahmad, Ali Al Bataineh +1 2026-09-19

    survey security

  26. Vision2CAD: A Visual Agent Harness for Explicit Geometry Referencing and Localization in Parametric CAD Modeling

    Xi Cheng, Chenxi Zhai, Hang Cheng +3 2026-09-19

    method harness

  27. Zero-Trust Authorization and Discovery for Enterprise MCP

    Huan Li, Yuwei Wang, Srinivasan Manoharan 2026-09-18

    method securityharnesstool use

  28. Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents

    Hafsa Akbar, Daniel Platnick, Marjan Alirezaie +1 2026-09-18

    method multi-agent incertain

  29. AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory

    Zijie Cao, Xijun Qu, Zhicheng Gu +7 2026-09-18

    method memory incertain

  30. An Agentic Just-in-Time Adaptive Intervention System for Personalized Sleep Support: Proof-of-Concept Study with N of 1 Data

    Nick Rezaee, Chelsea Boccagno 2026-09-18

    method

  31. CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents

    Tao Huang, Guosen Wu, Guolong Zheng +5 2026-09-18

    benchmark securityevaluation

  32. ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

    Qiang Zhang, Ruixue Ding, Fanrui Zhang +9 2026-09-18

    method

  33. Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

    Yining She, Lei Lin 2026-09-18

    study evaluation

  34. Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale

    Hao Fu, Baiting Zhu, Minglei Chen +2 2026-09-18

    method evaluationmulti-agentharness

  35. Two's a Crowd: Human and AI-Based Copresence for Developers with ADHD

    Veronica Pimenova, Seth Bernstein, Shalini Madan +2 2026-09-18

    study multi-agent incertain

  36. Can Agents Design Better Chips with a Higher Level Abstraction?

    Zijian Ding, Yang Zou, Yizhou Sun +1 2026-09-17

    method code released

  37. Chronicle: Cut-Point Replay for Regression Testing of LLM Agents

    Tisha Chawla, Susheem Koul 2026-09-17

    method evaluationcode releasedharness

  38. Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape

    Sarah Radway, Andrew Cheng, Vijay Janapa Reddi +1 2026-09-17

    study security incertain

  39. Language-model groups overstate consensus when replaying human deliberation on a reasoning task

    Tengfei Shao 2026-09-17

    study evaluationmulti-agentcode released

  40. How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

    Yukun Zhang, Kemu Xu, Yishen Chen 2026-09-17

    study harness

  41. A Proposal for an Agentic AI Architecture to Support Multi-Domain Decision-Making in the Brazilian Armed Forces

    Gioliano de Oliveira Braga, Sidnei Barbieri, Ágney Lopes Roth Ferraz +2 2026-09-17

    method

  42. Neuro-Symbolic Agentic AI for Networked Low-Altitude UAVs

    Yuqi Ping, Tianhao Liang, Nanchi Su +4 2026-09-17

    method

  43. A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents

    Haya Halimeh, Sascha Kaltenpoth, Kevin Bösch +1 2026-09-17

    study

  44. Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization

    Yingxuan Zhuang, Binhe Yu, Jingxiao Yang +6 2026-09-17

    method

  45. Rethinking Multi-Agent Collaboration: When More Is Less

    Yishuo Yuan, Yibo Wu, Yihan Zhang +5 2026-09-17

    method multi-agent

  46. AutoData: Agentic Search for Pre-training Data Selection

    Yan Meng, Dhruv Srikanth, Bingchen Zhao +2 2026-09-17

    method incertain

  47. SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes

    Mengxiao Wang, Nitesh Saxena 2026-09-17

    benchmark securityevaluation

  48. Self-Evolving Search Index

    Sangam Lee, Wonjae Lee, Sunghwan Kim +5 2026-09-17

    method incertain

  49. A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems

    Shaina Raza, Ahmed Y. Radwan, Imran Liaquat +1 2026-09-17

    benchmark securityevaluation incertain

  50. Closed-World Resolution Against Tool Hallucination in LLM Agents

    Laxmipriya Ganesh Iyer 2026-09-16

    benchmark securityevaluationcode releasedtool use

Les étiquettes sont attribuées par un classifieur, pas à la main, et leur exactitude n'a pas encore été mesurée. À utiliser pour restreindre la liste, pas comme un fait.