HAIC NeurIPS 2026
Second HAIC workshop · NeurIPS 2026

Human-AI Coevolution

Measuring human-agent teams in the agentic era.

  • Atlanta
  • December 12–13, 2026
  • In person, one day

Papers due September 2, 2026 (AoE) Extended from August 29 — no individual extensions Decisions September 29, 2026

About the workshop

Evaluation has not kept pace with deployment

Coding agents handle substantial portions of professional software workflows, clinical decision-support agents triage patients, and conversational agents mediate human learning and relationships. The community's evaluation methods, built largely for static benchmark performance on chat completion, have not adapted at the same pace.

A recent review of agentic-AI evaluation found that 83% of evaluations are dominated by technical metrics, while human-centered (30%), safety (53%), and economic dimensions (30%) remain peripheral (Jafari Meimandi et al., 2025). HAIC 2026 builds a methodological foundation for the empirical evaluation of human-agent teams. The central question is how to evaluate and govern human-agent systems rigorously as they coevolve with the people who use them, both in general deployment and in domains where the gap between benchmark and reality has the highest stakes?

This is the second edition of the HAIC series, following the inaugural ICLR 2025 workshop on human-AI coevolution. Where the first workshop mapped coevolution broadly across five themes, this one commits to a single focused operationalization: rigorous empirical evaluation of human-agent teams, with each theme anchored in recent peer-reviewed evidence.

The program grounds its discussion in case studies from high-stakes domains where the organizing team has direct research access: healthcare, mental health, aviation, and finance. The format weights discussion over talks, with breakouts feeding an open-problems registry and a community position paper.

Themes

How the themes relate

The themes are not independent. Deployment moves the validity target (Theme 1), the human feedback meant to correct course is itself contested (Theme 2), and evaluation must adapt as systems and users coevolve (Theme 3). Full descriptions are in the call for papers.

Theme 1

Validity of evaluation in deployed contexts

Benchmarks assume a static target. In deployment the target moves, and the constructs being measured, such as productivity, helpfulness, and safety, are themselves contested across domains. How do validity frameworks adapt?

Theme 2

Expert disagreement and the limits of human feedback

RLHF assumes aggregated feedback approximates a coherent target. A growing line of work treats disagreement as signal rather than noise. When does it mark evaluation invalidity, and when does it reflect domain pluralism that deployed systems should preserve?

Theme 3

Adaptive testing and continual evaluation

Deployed teams coevolve: skills reallocate, populations shift, distributions drift. A fixed test set can lose validity with no visible signal. We seek methods that track and respond to drift.

The three themes form a cycle: Theme 1 feeds Theme 2, Theme 2 feeds Theme 3, and Theme 3 returns to Theme 1 as drift moves the evaluation target again.

Key dates

Timeline

Dates follow the NeurIPS 2026 recommended workshop timeline. All deadlines are Anywhere on Earth. The submission deadline was extended from August 29 to September 2, 2026. This is the final date: the review process needs the remaining time, so no individual extensions can be granted beyond it.

Submission and notification schedule
MilestoneDateNotes
Submission deadline September 2, 2026 (AoE) Extended from August 29
Double-blind via OpenReview
Acceptance notification September 29, 2026 (AoE) Hard deadline
Camera-ready To be announced Set with authors after decisions
Workshop, Atlanta December 12–13, 2026 One day within that window, confirmed by NeurIPS
Invited

Speakers and panelists

Duncan Eddy

Duncan Eddy

Stanford University

Duncan Eddy is a research fellow at the Stanford Institute for Human-Centered AI and the Stanford Intelligent Systems Laboratory, funded by the Schmidt Sciences Foundation. He completed his PhD in aerospace engineering at Stanford, and his work builds robust decision making into autonomous robotic systems. His current research covers world modeling, automated red-teaming, and autonomous decision making for space systems, with active projects in cybersecurity, robotics, and spacecraft operations.

He was a founding engineer at Capella Space, the first US commercial synthetic-aperture radar imaging constellation, where he led the spacecraft operations group that built a fully automated constellation operations system. He went on to lead constellation operations and space safety at Project Kuiper, and was a Principal Applied Scientist at Amazon Web Services.

Melissa Valentine

Melissa Valentine

Stanford University

Melissa Valentine is a tenured professor of management science at Stanford University and a Senior Fellow at the Stanford Institute for Human-Centered Artificial Intelligence. She studies how technology is changing work and organizations, with particular attention to the evolving dynamics between AI systems and organizational design. She and her collaborators have received best paper awards at both management and computer science conferences, and she holds an NSF CAREER award.

Her work has been covered in the New York Times, the Wall Street Journal, Harvard Business Review, Wired, Fast Company, and the Financial Times. She holds a bachelor's degree from Stanford, a master's from NYU, and a PhD from Harvard.

Nataliya Kosmyna

Nataliya Kosmyna

MIT Media Lab

Nataliya Kosmyna is a Research Scientist in the Fluid Interfaces group at the MIT Media Lab and a Visiting Faculty Researcher at Google. She has more than sixteen years of experience designing and building end-to-end brain-computer interfaces, working across artificial intelligence, neuroscience, and human-computer interaction, and is interested in what a partnership between machine and human intelligence can look like.

Her research asks what the tools around us actually cost the people who use them, including her work on cognitive debt in LLM chatbot use, alongside cognitive alignment and the Humans Commons Licenses. Her systems and art projects have been deployed in classrooms, hospitals, workspaces, aerospace settings, low Earth orbit, and on the Moon. She received the L'Oréal-UNESCO Women in Science award in 2016.

Matthew Gombolay

Matthew Gombolay

Georgia Institute of Technology

Matthew Gombolay is the Stephen Fleming Early Career Associate Professor of Computing at the Georgia Institute of Technology. He holds a BS in mechanical engineering from Johns Hopkins University, and an SM in aeronautics and astronautics and a PhD in autonomous systems from MIT. Between his doctorate and joining the Georgia Tech faculty he was technical staff at MIT Lincoln Laboratory, where he transitioned his research to the US Navy and earned an R&D 100 Award.

His work has received best paper awards and nominations from the American Institute of Aeronautics and Astronautics, the ACM/IEEE Conference on Human-Robot Interaction, the Conference on Robot Learning, and Robotics: Science and Systems. He was selected as a DARPA Riser and holds a NASA Early Career Fellowship and an NSF CAREER award.

Liwei Jiang

Liwei Jiang

NVIDIA / UC Irvine (incoming)

Liwei Jiang is a Research Scientist at NVIDIA and an incoming Assistant Professor of Computer Science at the University of California, Irvine. She earned her PhD at the Paul G. Allen School of Computer Science and Engineering at the University of Washington. Her research addresses humanistic, pluralistic, and coevolutionary AI safety and alignment, spanning moral and pluralistic value reasoning in language models and data-, algorithm-, and system-level responses to socio-technical challenges in AI safety, security, and language model alignment.

Her work has received best paper awards at NeurIPS 2025, NAACL 2022, and CHI 2024, and outstanding paper awards at EMNLP 2023 and the AIA workshop at COLM 2025, and has been covered in the New York Times, Nature Outlook, IEEE Spectrum, and Wired. She co-organizes the MP2 and SoLaR workshops and co-leads the Guardrails and Security for LLMs tutorial at ACL 2025.

Organizing committee

Organizers

Call for papers

Scope of submissions

We invite new evaluation methods and benchmarks, datasets of human-agent interaction, reproducible benchmark critiques, case studies from high-stakes domains, and position papers. The workshop is non-archival, and each submission receives three double-blind reviews via OpenReview.