BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20251016T135131Z
LOCATION:Grand Hall G/H
DTSTART;TZID=America/Chicago:20251015T152000
DTEND;TZID=America/Chicago:20251015T154000
UID:HFESAM_ASPIRE 2025_sess170_LBR137@linklings.com
SUMMARY:Holistic Evaluation of Explainable Artificial Intelligence for Hum
 an-Autonomy Teaming
DESCRIPTION:Mark Boyer (University of Colorado Boulder, U.S. Air Force) an
 d Torin Clark (University of Colorado Boulder)\n\nHuman-autonomy teaming w
 ith an Explainable AI (XAI) system is being proposed in a variety of high-
 consequence, dynamic settings such as military aviation or space explorati
 on.  Rigorous, human-centered evaluation of XAI systems is still a nascent
  field and often lacks the full range of human factors evaluation. We cond
 ucted an evaluation of the effect explanation type for AI route planning o
 n team performance, mental workload, trust, situation awareness (SA), and 
 user preference in a realistic space exploration simulator. Participants (
 N=16, 10M/6F) simultaneously drove a simulated Mars rover while supervisin
 g and directing a highly-autonomous unmanned rover.  Participants received
  various forms of explanations of AI-generated routes, including global go
 al explanation, contrastive explanation, or deductive explanation.  Linear
  Mixed Effects models were used to account for trial number and condition 
 relative to all outcomes measured, with a baseline condition of “AI agent 
 available, no explanation provided”. Performance on the manual driving tas
 k was better with global explanations (p=0.02). Performance on the autonom
 y supervision task was best with all explanations (p=0.01). Performance on
  the combined task was best with global and contrastive explanations (p=0.
 016) or all explanations (p=0.021), but worse when no AI agent was availab
 le (p=0.02). Explanation type also led to changes in workload ratings (p=0
 .0.18), explainability ratings (p<0.001), trust ratings (p<0.001), and usa
 bility ratings (p=0.005). SAGAT performance did not change with explanatio
 n type (p=0.24).  These results begin to establish a systematic means of a
 ssessing  explanation type through a human-centered evaluation of XAI syst
 ems.\n\nTrack: General Sessions\n\n
END:VEVENT
END:VCALENDAR
