Skip to content
Field NoteOperationsAugust 31, 2026 · 7 min read

Drone Autonomy and Workload: What a 20-Person Field Study Can Support

A bounded reading of NASA-TLX evidence comparing two drone inspection methods across a hangar and greenhouse study.

A lower workload score can justify a larger trial. It cannot, by itself, justify a claim of safer operations.

Automation is frequently sold with a broad promise: less work for the operator. Human-factors evidence demands a more precise question. Which tasks were compared, with which participants, in what environment, and using which measure? Hirad Goudarzi’s 2024 University of Bristol doctoral dissertation provides a useful bounded case. Twenty participants completed drone inspection activities using two methods identified in the thesis as CMR and WMA. CMR means Conventional Waypoint Following with Manual Recovery: the participant manually steers away after detecting a hazard. WMA means Waypoint Following with Movement Authorization: the participant authorizes or withholds the next movement, and a withheld authorization sends the aircraft to a designated safe location. Their weighted NASA Task Load Index scores averaged 38.68 for CMR and 32.97 for WMA.

The arithmetic difference is 5.71 points on the reported 0–100 workload scale. That subtraction is calculated from the two reported means; it is not an additional effect estimate quoted from the dissertation. A paired t-test reported (t=2.76) and (p=.012), indicating that the observed within-participant difference was statistically distinguishable under the study design. The result supports a claim about subjective workload in this experiment. It does not establish fewer accidents, better mission outcomes, or a universal advantage for drone autonomy.

Drone autonomy and operator workload evidence map

Editorial evidence graphic. Scores are subjective weighted NASA-TLX results from 20 participants in two inspection settings, not an operational accident-rate study.

How the field comparison was run

The study involved 20 people with varied drone experience. The first 15 performed an inspection task in a hangar at the Bristol Robotics Laboratory’s Severe Accident Centre setting described in the thesis. The remaining five performed a greenhouse inspection at Fenswood Farm. Participants experienced both methods, and the study counterbalanced their order to reduce the risk that practice or fatigue would systematically favor one condition.

Both conditions used the same waypoint mission. The difference was how the participant responded to a person entering the operating area, the only contingency included in this experiment. A safety pilot retained ultimate control through a separate primary radio link. Battery, distance, and navigation-failure contingencies were not presented to participants. WMA should therefore be read as a supervised, Ground Control Automaton (GCA)-enabled interaction method—not as fully uncrewed operation.

After each flight, participants completed a weighted NASA-TLX assessment and a three-dimensional Situation Awareness Rating Technique assessment. The WMA condition also received a System Usability Scale questionnaire. NASA-TLX is a structured subjective workload instrument. It organizes perceived demands such as mental, physical, and temporal demand, effort, performance, and frustration. It is valuable because operator burden is not fully visible in flight time or path accuracy. It remains self-report evidence rather than a direct physiological measure or accident outcome.

The statistical checks reported for the workload comparison included Shapiro–Wilk tests: CMR (W=.911, p=.067) and WMA (W=.905, p=.052). The thesis then reports a paired t-test for the two conditions. Pairing is appropriate to the study question because the same participant contributes a score under each method, so the comparison focuses on within-person change instead of treating the two sets as unrelated groups.

The results, with boundaries attached

MeasureCMRWMAComparison
Participants2020Within-participant, counterbalanced comparison
Weighted NASA-TLX mean38.6832.97Reported 0–100 workload score
Standard deviation10.068.75Variation among participant scores
Mean-score difference5.71 points, calculated as 38.68 − 32.97
Paired test(t=2.76, p=.012)

The thesis also describes lower temporal demand under WMA and records that some participants felt more rushed under CMR. That detail helps explain the total-score difference, but it should not be converted into a general rule that automation always reduces time pressure. The two environments and tasks shape the demand profile. The small greenhouse subgroup also means that the study is not designed to provide a stable hangar-versus-greenhouse estimate.

An average can hide meaningful individual differences. The dissertation notes that some experienced participants found or preferred the conventional method more easily. That is operationally important. Automation can reduce routine control burden while introducing supervision, mode awareness, recovery, or trust-calibration demands. A workforce with extensive manual skill may experience a transition differently from novice operators. Procurement should therefore examine the score distribution, exceptions, and failure-recovery behavior rather than purchase from the mean alone.

What may and may not be claimed

A faithful sentence is: “In a 20-participant, counterbalanced drone-inspection study spanning a hangar and greenhouse setting, mean weighted NASA-TLX was 38.68 under CMR and 32.97 under WMA; the dissertation reported a paired-test result of (p=.012).” This gives readers the population, context, measure, values, and test.

The study does not support “autonomous drones are 15% safer” or “automation reduces operator workload by 5.71%.” The 5.71 value is a point difference, not a percentage. Dividing it by the CMR mean would produce a relative arithmetic change, but that derived percentage could imply a precision and transferability the study did not establish. More importantly, lower subjective workload is not identical to higher situation awareness, better detection performance, or lower incident probability.

Nor should statistical significance be presented as operational importance. The (p)-value addresses compatibility with a no-difference model under assumptions; it does not quantify the probability that WMA is better, the size of benefit in a new operation, or the cost of implementation. Decision-makers still need effect uncertainty, task performance, intervention failures, training time, recovery actions, and mission consequences.

Turning a dissertation result into a local test

A drone operator or buyer can use this result to design a focused evaluation. Select one recurring inspection with measurable completion criteria. Use a within-participant crossover where practical, counterbalance method order, and document previous experience. Record task completion, missed inspection points, interventions, recoveries, flight time, and NASA-TLX after each condition. Add a situation-awareness measure and a short usability instrument only if the team has a plan for interpreting them.

Separate normal operation from exceptions. An automated path may look favorable during nominal flight but impose high workload when positioning fails, communications degrade, or a human must resume control. Include a controlled, safe recovery scenario and record the time to recognize, decide, and stabilize. This does not require creating a hazardous event; simulation or a bounded test environment can expose the transition burden.

Define the decision threshold before the trial. A lower average NASA-TLX score should not pass a method that misses critical inspection evidence. Conversely, a modest workload change may still matter if it reduces temporal pressure during a high-consequence task. Workload belongs in a multi-metric acceptance case, not at the top of a single-metric leaderboard.

Evidence-use checklist

  • Name the two compared methods exactly as defined by the source or local protocol.
  • State that the dissertation study included 20 participants.
  • Preserve the 15-person hangar and five-person greenhouse context.
  • Report NASA-TLX as subjective workload on the stated scale.
  • Keep means and standard deviations together.
  • Label 5.71 as an arithmetic point difference, not a quoted percentage effect.
  • Include the paired design and (t=2.76, p=.012) when discussing inference.
  • Do not convert workload into accident reduction or safety certification.
  • Inspect experienced-operator exceptions and recovery workload.
  • Require task quality and situation-awareness evidence beside workload.
  • Hirad Goudarzi (2024), In Praise of Lights and Clockwork: A Simple Approach to Drone Autonomy, doctoral dissertation, University of Bristol: official university record and full text. The method definitions and safety-pilot setup are in Chapter 4, Sections 4.5.1–4.5.2; workload results are in Section 4.6 and Figures 4.8–4.10.
  • Reproducible discovery query: Google Scholar exact-title search. Scholar indexing can change; the university record is the controlling bibliographic source used here.

This field note paraphrases the dissertation. It does not reproduce thesis prose and should not replace review of the full methods, figures, and limitations before a procurement or safety decision.

Tags
drone autonomyoperator workloadNASA-TLXhuman factorsdrone inspection
More in Operations
Reviewed insight feed

Follow evidence-reviewed field notes.

Subscribe to the RSS feed for reviewed articles on drone operations, bird-strike risk, CBRN readiness, and aerospace ESG. We do not collect an email address until a verified mailing service is available.