About The Workshop
AI agents are shifting from static pretrained systems to adaptive systems that update their models, memories, tools, and architectures over time. Evolution can arise through self-evolution, co-evolution with other agents or users, and environment-driven adaptation under non-stationary tools, interfaces, and tasks. As agents operate in long-horizon workflows, the central question extends beyond single-turn performance to whether that evolution remains trustworthy: small errors can accumulate, persist, and amplify in stateful systems.
The Core Challenge: Many trustworthiness failures in evolving agents are systems failures. Corrupted memory persists when there is no provenance, observation–inference separation, or rollback. Alignment drift accumulates when biased updates are applied irreversibly, without staging, regression checks, or an audit trail. Unsafe tool use persists when tool calls lack permission checks, sandboxing, and revocation. In each case, the model is asked to enforce by disposition what mature computing systems enforce by construction.
This workshop examines trustworthy evolution through the lens of an agent operating system: the memory, capability, runtime, update-management, and observability subsystems that sit beneath an evolving agent. Operating systems did not make programs correct; they made unreliable programs safer to run through isolation, privilege separation, recoverable state, versioning, and auditing. We ask which classical OS abstractions transfer to stochastic, learning components and which break.
The central question is: how should we design the systems layer beneath an evolving agent so that its evolution remains reliable over long horizons? We bring together agent learning, ML systems, and AI safety to study how agents change, how errors accumulate, and how harmful updates can be detected, attributed, and reversed without sacrificing performance.
Topics
Each topic pairs a dimension along which agents evolve with the subsystem that must keep that evolution trustworthy.
Provenance and write policies, consolidation and forgetting, transactional writes and snapshots, isolation across users and tenants, hallucination propagation, and leakage from persistent state.
Tools as capability-mediated system calls: least privilege, separation of untrusted content from privileged execution, sandboxing and verification of synthesized tools, revocation, and supply-chain risk in tool registries.
Self-improvement as a release process: versioning of prompts, policies, and weights; staged rollout and canarying; regression suites; rollback without collateral capability loss; reliability of self-generated critiques; and prevention of reward hacking.
Lifecycle and concurrency for agent populations, resource accounting over tokens, time, and side effects, preemption of runaway loops, checkpointing, blast-radius limits, and emergent failures under co-adaptation.
Structured traces attributing an outcome to the update that caused it, continuous monitoring rather than one-off benchmarking, oversight interfaces that escalate before irreversible updates, and stateful long-horizon testbeds.
Call For Papers
The AAAI 2027 Workshop on Trustworthy Evolution of AI Agents invites submissions that treat trustworthy agent evolution as a systems problem. We welcome mechanisms and abstractions, system designs, benchmarks, negative results, and position papers.
Key Dates
- Submission Deadline: November 20, 2026, AoE
- Notification Date: December 2, 2026, AoE
- Workshop Date: February 22–23, 2027 (one-day workshop; exact day TBA)
Deadlines are strict and will not be extended. All deadlines follow the Anywhere on Earth (AoE) timezone.
Submission Site
Submissions will be managed via OpenReview.
Link to OpenReview Submission Portal (TBD)
Submission Guidelines
Formatting Requirements
Submissions must be in English and follow the AAAI 2027 author kit / workshop LaTeX template (TBD).
Papers must be submitted as a single PDF file. We welcome:
- Full Papers (8 pages)
- Short Papers / Position Papers (4 pages)
- References and appendices are usually not included in the page limit.
Anonymity
The workshop follows a double-blind review process. Submissions must be anonymized.
Contact
For questions, please contact at trustworthyagentevolving@gmail.com.
Speakers and Panelists
Schedule
Tentative workshop schedule. All talks include a Q&A session.
| Time (EST) | Session | Speaker | Talk Title |
|---|---|---|---|
| 08:00 – 08:15 | Opening remarks: Trustworthy Agent Evolution as a Systems Problem | Organizers | |
| 08:15 – 09:00 | Keynote 1 (Foundations) | Prof. Yoshua Bengio (Mila, Université de Montréal, LawZero) | Safeguarding Agent Evolution: The Scientist AI Approach to Catastrophic Risk |
| 09:00 – 09:45 | Spotlight orals | 3 orals (prioritize junior and underrepresented researchers) | |
| 09:45 – 10:30 | Keynote 2 (Capability & Isolation) | Prof. Dawn Song (University of California, Berkeley) | Privilege, Isolation, and Verifiable Execution in Agent Runtimes |
| 10:30 – 12:00 | Poster session 1 | Poster session, lunch break, and Q&A round table for junior researchers | |
| 12:00 – 12:45 | Keynote 3 (Runtime & Infrastructure) | TBD (tentative) | What an Operating System for Agents Should Actually Provide |
| 12:45 – 13:15 | Spotlight orals | 2 orals (prioritize junior and underrepresented researchers) | |
| 13:15 – 14:00 | Keynote 4 (Measurement) | Prof. Sanmi Koyejo (Stanford University) | From Benchmarks to Continuous System-Level AI Measurement |
| 14:00 – 15:30 | Poster session 2 | Poster session and coffee break | |
| 15:30 – 16:45 | Panel discussion (Slido Q&A) | Prof. Siva Reddy (Mila, McGill); TBD (tentative) | See panel topics below |
| 16:45 – 17:00 | Closing remarks and Outstanding Paper Awards | Organizers | |
Talk titles are tentative and will be finalized with the speakers. Panel topics:
- Which OS abstractions survive contact with stochastic agents?
- Rollback and reversibility: can a self-improving agent be un-improved?
- Where should trust be enforced: in the model, or in the substrate?
- Human–agent co-evolution: oversight as a system interface.
- Deployment reality: what infrastructure do production agents actually need?
Student Registration Grant
The Trustworthy Evolution of AI Agents Workshop is pleased to announce limited financial support for student authors of accepted workshop papers.
Thanks to our generous sponsors, we are offering registration reimbursement grants to ensure broader student participation in the community.
💡 Eligibility & Coverage
- Only the student registration fee for workshop is eligible for reimbursement.
- Support will be issued via reimbursement after the conference (not as direct payment).
- Travel, accommodation, or other expenses are not covered under this program.
💬 Reimbursement Policy
Reimbursements will be processed after the conference upon submission of a valid payment receipt for the student registration.
This ensures funds are distributed fairly to participants who attend and present at the workshop.
🗓️ Key Dates
- Application Deadline: TBD
- Notification of Results: TBD
🎯 Priority Considerations
Preference will be given to:
- students presenting accepted papers at our Trustworthy Evolution of AI Agents workshop.
- students without other funding support (e.g., from advisor or institution).
- students from underrepresented regions or institutions.
Organizers
This workshop is organized by
Supporting Team
Beyond the organizers, the following members support workshop delivery.
Advisory Board
This workshop is advised by
