AAAI 2027 Workshop on Trustworthy Evolution of AI Agents

An Agent Operating System Perspective
February 22–23, 2027
Palais des congrès de Montréal, Montréal, Canada

About The Workshop

AI agents are shifting from static pretrained systems to adaptive systems that update their models, memories, tools, and architectures over time. Evolution can arise through self-evolution, co-evolution with other agents or users, and environment-driven adaptation under non-stationary tools, interfaces, and tasks. As agents operate in long-horizon workflows, the central question extends beyond single-turn performance to whether that evolution remains trustworthy: small errors can accumulate, persist, and amplify in stateful systems.

The Core Challenge: Many trustworthiness failures in evolving agents are systems failures. Corrupted memory persists when there is no provenance, observation–inference separation, or rollback. Alignment drift accumulates when biased updates are applied irreversibly, without staging, regression checks, or an audit trail. Unsafe tool use persists when tool calls lack permission checks, sandboxing, and revocation. In each case, the model is asked to enforce by disposition what mature computing systems enforce by construction.

This workshop examines trustworthy evolution through the lens of an agent operating system: the memory, capability, runtime, update-management, and observability subsystems that sit beneath an evolving agent. Operating systems did not make programs correct; they made unreliable programs safer to run through isolation, privilege separation, recoverable state, versioning, and auditing. We ask which classical OS abstractions transfer to stochastic, learning components and which break.

The central question is: how should we design the systems layer beneath an evolving agent so that its evolution remains reliable over long horizons? We bring together agent learning, ML systems, and AI safety to study how agents change, how errors accumulate, and how harmful updates can be detected, attributed, and reversed without sacrificing performance.



Topics

Each topic pairs a dimension along which agents evolve with the subsystem that must keep that evolution trustworthy.

(1) Trustworthy Memory Evolution (the state subsystem):

Provenance and write policies, consolidation and forgetting, transactional writes and snapshots, isolation across users and tenants, hallucination propagation, and leakage from persistent state.

(2) Trustworthy Tool and Skill Evolution (the capability interface):

Tools as capability-mediated system calls: least privilege, separation of untrusted content from privileged execution, sandboxing and verification of synthesized tools, revocation, and supply-chain risk in tool registries.

(3) Trustworthy Model and Policy Evolution (the update and release layer):

Self-improvement as a release process: versioning of prompts, policies, and weights; staged rollout and canarying; regression suites; rollback without collateral capability loss; reliability of self-generated critiques; and prevention of reward hacking.

(4) Trustworthy Multi-Agent Co-Evolution (the runtime and scheduler):

Lifecycle and concurrency for agent populations, resource accounting over tokens, time, and side effects, preemption of runaway loops, checkpointing, blast-radius limits, and emergent failures under co-adaptation.

(5) Evaluating Trustworthy Evolution (the observability and audit layer):

Structured traces attributing an outcome to the update that caused it, continuous monitoring rather than one-off benchmarking, oversight interfaces that escalate before irreversible updates, and stateful long-horizon testbeds.



Call For Papers

The AAAI 2027 Workshop on Trustworthy Evolution of AI Agents invites submissions that treat trustworthy agent evolution as a systems problem. We welcome mechanisms and abstractions, system designs, benchmarks, negative results, and position papers.

Key Dates

  • Submission Deadline: November 20, 2026, AoE
  • Notification Date: December 2, 2026, AoE
  • Workshop Date: February 22–23, 2027 (one-day workshop; exact day TBA)

Deadlines are strict and will not be extended. All deadlines follow the Anywhere on Earth (AoE) timezone.


Submission Site

Submissions will be managed via OpenReview.

Link to OpenReview Submission Portal (TBD)


Submission Guidelines

Formatting Requirements

Submissions must be in English and follow the AAAI 2027 author kit / workshop LaTeX template (TBD).

Papers must be submitted as a single PDF file. We welcome:

  • Full Papers (8 pages)
  • Short Papers / Position Papers (4 pages)
  • References and appendices are usually not included in the page limit.
Anonymity

The workshop follows a double-blind review process. Submissions must be anonymized.

Contact

For questions, please contact at trustworthyagentevolving@gmail.com.

Speakers and Panelists

Yoshua Bengio
Yoshua Bengio

Mila, Université de Montréal & LawZero

Dawn Song
Dawn Song

University of California, Berkeley

Sanmi Koyejo
Sanmi Koyejo

Stanford University

Siva Reddy
Siva Reddy

Mila & McGill University

Schedule

Tentative workshop schedule. All talks include a Q&A session.

Time (EST) Session Speaker Talk Title
08:00 – 08:15 Opening remarks: Trustworthy Agent Evolution as a Systems Problem Organizers
08:15 – 09:00 Keynote 1 (Foundations) Prof. Yoshua Bengio (Mila, Université de Montréal, LawZero) Safeguarding Agent Evolution: The Scientist AI Approach to Catastrophic Risk
09:00 – 09:45 Spotlight orals 3 orals (prioritize junior and underrepresented researchers)
09:45 – 10:30 Keynote 2 (Capability & Isolation) Prof. Dawn Song (University of California, Berkeley) Privilege, Isolation, and Verifiable Execution in Agent Runtimes
10:30 – 12:00 Poster session 1 Poster session, lunch break, and Q&A round table for junior researchers
12:00 – 12:45 Keynote 3 (Runtime & Infrastructure) TBD (tentative) What an Operating System for Agents Should Actually Provide
12:45 – 13:15 Spotlight orals 2 orals (prioritize junior and underrepresented researchers)
13:15 – 14:00 Keynote 4 (Measurement) Prof. Sanmi Koyejo (Stanford University) From Benchmarks to Continuous System-Level AI Measurement
14:00 – 15:30 Poster session 2 Poster session and coffee break
15:30 – 16:45 Panel discussion (Slido Q&A) Prof. Siva Reddy (Mila, McGill); TBD (tentative) See panel topics below
16:45 – 17:00 Closing remarks and Outstanding Paper Awards Organizers

Talk titles are tentative and will be finalized with the speakers. Panel topics:

  • Which OS abstractions survive contact with stochastic agents?
  • Rollback and reversibility: can a self-improving agent be un-improved?
  • Where should trust be enforced: in the model, or in the substrate?
  • Human–agent co-evolution: oversight as a system interface.
  • Deployment reality: what infrastructure do production agents actually need?

Student Registration Grant

The Trustworthy Evolution of AI Agents Workshop is pleased to announce limited financial support for student authors of accepted workshop papers.

Thanks to our generous sponsors, we are offering registration reimbursement grants to ensure broader student participation in the community.

💡 Eligibility & Coverage

  • Only the student registration fee for workshop is eligible for reimbursement.
  • Support will be issued via reimbursement after the conference (not as direct payment).
  • Travel, accommodation, or other expenses are not covered under this program.

💬 Reimbursement Policy

Reimbursements will be processed after the conference upon submission of a valid payment receipt for the student registration.

This ensures funds are distributed fairly to participants who attend and present at the workshop.

🗓️ Key Dates

  • Application Deadline: TBD
  • Notification of Results: TBD

🎯 Priority Considerations

Preference will be given to:

  • students presenting accepted papers at our Trustworthy Evolution of AI Agents workshop.
  • students without other funding support (e.g., from advisor or institution).
  • students from underrepresented regions or institutions.

Organizers

This workshop is organized by

Haolun Wu
Haolun Wu

Stanford University & Mila
Primary Contact

Negar Arabzadeh
Negar Arabzadeh

University of California, Berkeley

Wenyue Hua
Wenyue Hua

Microsoft Research

Bang Liu
Bang Liu

Mila & Université de Montréal

Laurent Charlin
Laurent Charlin

Mila & HEC Montréal

Supporting Team

Beyond the organizers, the following members support workshop delivery.

Ye Yuan
Ye Yuan

Mila & McGill University
Web and Technical Chair

Yankai Chen
Yankai Chen

MBZUAI
Review and Program Coordinator

Tianyu Shi
Tianyu Shi

McGill University
Publicity and Sponsorship Chair

Advisory Board

This workshop is advised by

Preslav Nakov
Preslav Nakov

MBZUAI

Xue (Steve) Liu
Xue (Steve) Liu

MBZUAI & McGill University