Source: https://911planesresearch.substack.com/p/why-the-manual-or-automated-study
What if the central experiment in a study claiming to distinguish “manual” from “automated” flight never actually tested the aircraft it was supposed to investigate?In May 2026, The International Center For 9/11 Justice published a study it conducted called Manual or Automated? A Flight Simulation Study and Analysis of Reported Aircraft Maneuvers on September 11, 2001, Which was authored by Dr. Piers Robinson. Manual or Automated? study asks whether the reported September 11th flight paths could have been manually flown—but its principal simulator experiment used a Boeing 737 to investigate manoeuvres performed by Boeing 767s and a Boeing 757.The study is presented as an empirical investigation into whether the aircraft reported to have struck the World Trade Center and Pentagon were manually controlled or operated using some form of automated guidance. Its methodology combines reported flight-path data, three-dimensional modelling and a series of flight-simulator experiments involving seven pilots. The study ultimately concludes that the reported flight paths are more consistent with automated control than manual piloting.There is however, a fundamental methodological problem at the heart of the study. The experiment does not actually test the proposition it claims to test in a sufficiently controlled or aircraft-specific manner.The problem is particularly serious because the study uses a Boeing 737 simulator to investigate the performance of Boeing 767 and Boeing 757 aircraft. The authors themselves acknowledge that the 737simulator may have been less manoeuvrable than the aircraft whose reported flight paths they were attempting to reproduce. That admission is not a minor technical qualification. It goes directly to the validity of the experiment.The Central Experimental Mismatch The photograph below, highlights the difference in the size between a 737 aircraft and 767 aircraft. The objective is to give the reader of this article an idea of the difference between the two aircrafts, which is integral to understanding the veracity of the analysis I have conducted. Below, is the comparison specifications of a BOEING 767-200/200ER vs BOEING 737-800, to address Aidan Monaghan’s comment disputing the Delta Airlines aircraft photograph comparison I used above. As you can observe, the 737, is still considerably smaller than a 767. A 737 simulator was used to study 767 and 757 flight paths. The study states explicitly that, after searching for commercially available simulators, the researchers obtained a full-motion Boeing 737 simulator. The aircraft involved in the principal simulations, however, included Boeing 767s and a Boeing 757. United Airlines 175 (UAL 175) and American Airlines 11 (AAL 11) are identified by the study as Boeing 767s, while American Airlines 77 (AAL 77) is identified as a Boeing 757.This creates an immediate external-validity problem. A simulator is not merely a visual representation of an aircraft. Its aerodynamic model, control response, mass and inertia characteristics, engine performance, control-system behaviour and flight-envelope modelling determine how the simulated aircraft responds to pilot inputs.Consequently, demonstrating that a pilot could or could not perform a particular manoeuvre in a 737 does not, by itself, establish that the same pilot could or could not perform that manoeuvre in a 767 or 757.The study attempts to overcome this problem by having two highly experienced pilots evaluate the simulator. These pilots did not actually fly the 9/11 manoeuvres during the preliminary evaluation. Instead, they evaluated the simulator’s handling at high and low speeds and concluded that it provided a “reasonable approximation” of how a real 767 or 757 would handle.That is not the same thing as demonstrating equivalence. The critical question is not whether two experienced pilots considered the simulator generally representative, the question is whether the simulator reproduced the specific performance characteristics relevant to the manoeuvres being tested. Those characteristics include roll rate, turn response, pitch response, load-factor response, energy retention, control sensitivity, high-speed behaviour and the ability to execute the precise manoeuvres being investigated.Indeed, the study itself subsequently acknowledges that the Boeing 737 simulator may have been “less manoeuvrable” than a real Boeing 767 or 757. It specifically states that this could have made the final SouthTower turn and Pentagon pull-out more difficult than they would have been in the actual aircraft. This concession is potentially devastating to the experiment’s principal inference.If the experimental apparatus systematically makes the manoeuvre harder than it was in the aircraft under investigation, then a low completion rate cannot safely be interpreted as evidence that the real-world manoeuvre was inherently difficult for the actual aircraft. At most, it establishes that the manoeuvre was difficult in that simulator, under those experimental conditions.The study acknowledges the simulator may have biased the resultsThis is particularly important because the study is not merely using the simulator to demonstrate what a manoeuvre looks like, it uses pilot success and failure rates to support a substantive conclusion about the likelihood of “manual” control.The South Tower first-attempt results are reported as zero successful completions among five relevant attempts, while the Pentagon experiment similarly produced no first-attempt completions. The study then uses these results as evidence that the manoeuvres were inherently difficult and therefore unlikely to have been performed manually.The report simultaneously acknowledges several factors capable of influencing those failure rates:Uncertainty over the correct starting offset.Uncertainty over precisely when pilots should initiate the turn.Potentially inferior manoeuvrability of the 737 simulator.Limitations in the simulator’s graphics.Difficulties in determining appropriate instructions; and the possibility that different simulators and instructions could produce higher completion rates.These are not peripheral issues, they are experimental confounders. If the independent variable is supposed to be the difficulty of the real-world aircraft manoeuvre, but the result is also strongly influenced by simulator fidelity, instructions, starting position, timing and visual presentation, then the experiment does not isolate the variable it claims to measure.The study therefore has a classic identification problem, it cannot determine how much of the observed difficulty arose from the manoeuvre itself and how much arose from the experimental apparatus.The experiment does not reproduce the actual eventThere is a second fundamental problem. The study asks pilots to reproduce a reconstructed flight path, but the pilots are not simply placed in an aircraft and told to fly toward the target using the information available to a pilot at the time, they are given explicit experimental instructions.For the South Tower experiment, pilots were instructed to maintain an offset and then, at approximately ten seconds before impact, were told to initiate the turn. For the second stage, the instructions were refined further by specifying a particular magnetic heading of 063 degrees to ensure the required offset.This creates an important distinction between, “Can a pilot independently reproduce the observed flight path?” and “Can a pilot execute a manoeuvre after being explicitly instructed when and how to initiate it?”Neither experiment perfectly reproduces the circumstances of the original event. The alleged hijackers did not receive an instruction from a researcher saying, in effect, “turn now.” Nor do we know precisely what visual references, navigation information, headings, targets or mental model were available to the person controlling the aircraft.The experiment therefore cannot legitimately treat pilot failure under its artificial instructions as a direct measurement of the probability that the original aircraft could have been manually controlled.The South Tower experiment was not even a complete replication of the reported manoeuvre. This is especially significant, because the study describes the reported UAL 175 manoeuvre as involving a rapid descent followed by a pull-out and banked turn immediately before impact. Yet the researchers explicitly state that, to avoid over-complicating the experiment, they did not ask the pilots to reproduce the final pull-out, They concentrated on the final turn.That means the experiment did not reproduce the complete manoeuvre it was supposedly evaluating. This matters because the study’s broader argument depends upon the supposed combination of descent, pull-out, bank and final alignment. If only one component is tested, the resulting experiment cannot establish the difficulty of the entire sequence, so the conclusion therefore becomes stronger than the experiment warrants.Failure to reproduce, does not establish “automation”Perhaps the most important logical flaw is the leap from difficulty to mechanism. Suppose, for the sake of argument, that the study successfully demonstrated that the reported manoeuvre was exceptionally difficult for a low-experience pilot. That would establish only that the manoeuvre was difficult, it would not establish that an automated guidance system performed it. There are numerous alternative explanations greater-than-assumed pilot skillPrior experience with large aircraft.Prior simulator practice.Prior knowledge of the intended route.Different aircraft performance.Different control inputs.Different starting conditions.Different interpretation of the radar-derived flight path.Errors or uncertainties in the reconstructed trajectory.Different visual conditions.Deliberate rather than intuitive piloting.A combination of manual and automated control.Use of navigation equipment without autonomous flight, or simply a successful manual manoeuvre with low probability.The experiment does not experimentally eliminate these alternatives, moreover, the study itself acknowledges that its first-attempt results are uncertain and that different instructions and different simulated aircraft could produce higher completion rates.Consequently, the correct inference from the experiment is potentially, “The manoeuvre proved difficult for these pilots under these experimental conditions.”The study instead moves toward, “The manoeuvre was unlikely to have been manually flown and is better explained byautomated control.”That is a substantially stronger proposition, the latter does not follow logically from the former without additional evidence.The sample is extremely smallThe study involved seven pilots, divided into high- and lower-experience groups. The lower-experience group consisted of four pilots with approximately 2,000, 150, 300 and 500 hours respectively. This is a very small experimental sample. Small samples can be perfectly legitimate in exploratory aviation research, particularly where simulator time is expensive. But they cannot support sweeping probabilistic claims without considerable caution. For example, if five people fail a manoeuvre on their first attempt, that does not establish that a particular class of pilot has a high probability of failing the manoeuvre in the real world.The study provides no sufficiently developed statistical model translating simulator performance into the probability of real-world success, therefore there is a substantial gap between the observed experimental outcome and the population-level conclusion.The “first attempt” assumption is problematicThe study places considerable importance on first attempts. Its reasoning is essentially that because the aircraft supposedly struck their targets on the first attempt, the simulator pilots should also be evaluated on their first attempts. The study therefore deliberately focuses on first-attempt performance.But this creates an important assumption, that the original pilots had no prior relevant experience with the aircraft, route, manoeuvre or circumstances. The study argues that specific practice of the exact manoeuvres is implausible. That is an assertion, not an experimentally established fact. Moreover, general experience can matter enormously. A pilot does not need to have rehearsed an exact trajectory to possess transferable skills involving aircraft control, high-speed turns, energy management, visual targeting and navigation. The study’s own pilots demonstrate this problem. Pilot 6 explicitly attributed some of his difficulty to unfamiliarity with the 737 simulator and stated that, compared with an aircraft he normally flew, he believed he could adapt the turn more sharply.That statement directly undermines any simplistic interpretation of simulator failure as evidence of the intrinsic impossibility of the manoeuvre.The pilot was learning the behaviour of the simulator. That is precisely the problem with using first-attempt simulator performance as a proxy for first attempt performance in a different aircraft.The experiment contains a built-in learning effectThe report itself demonstrates that pilots became more successful with repeated attempts. For the Pentagon manoeuvre, the study records multiple successful completions on later runs by experienced pilots, despite failures on earlier attempts.This shows that simulator familiarity substantially affected performance. That observation is important because it means the experiment is simultaneously measuring at least two things:The difficulty of the manoeuvreThe pilot’s learning curve on the simulator.The study treats the first attempt as particularly probative, but the learning effect makes it difficult to separate the two.A scientifically stronger experiment would have required pilots to be thoroughly familiarised with the exact aircraft model and simulator controls before their test runs, while keeping the specific manoeuvre itself unknown. That would isolate aircraft-control skill from simulator unfamiliarity. Instead, the study allows only at least 15 minutes of simulator familiarisation before testing.For a pilot being asked to perform an extreme manoeuvre at very high speed in an aircraft type different from the one they normally fly, that is hardly a robust control for simulator familiarisation.The study relies heavily on reconstructed flight pathsAnother fundamental issue concerns the dependent variable itself. The researchers are not measuring the original control inputs. For UAL 175 and AAL11, the study acknowledges that no FDR was retrieved. Consequently, the researchers reconstruct aspects of the aircraft’s flight behaviour from radar and other evidence. But a reconstructed trajectory is not equivalent to knowing the actual control inputs.A radar-derived track can provide information about position, altitude, speed and heading, but it does not directly tell us whether a particular movement was generated by, manual control, autopilot, flight-management-system commands, flight-director guidance, some combination of systems, or another control mechanism.The study therefore begins with an observational reconstruction and then uses that reconstruction as the target for a simulator experiment. That introduces uncertainty at both ends of the analysis. Uncertainty in the reconstructed original trajectory + uncertainty in the simulator’s ability to reproduce the aircraft = uncertainty in the experimental conclusion. Those uncertainties should compound rather than disappear.The study itself identifies unresolved questions about the underlying dataAn especially revealing feature of the report is that its final sections acknowledge that further examination of the flight-path and FDR evidence is still necessary.The authors themselves say that fuller examination of earlier flight phases is warranted, including questions surrounding AAL 77’s delayed U-turn and loss of radar tracking. They also state that further examination of flight and airport logs, including ACARS data, is needed.This creates a methodological tensionIf important elements of the underlying flight-path evidence remain unresolved, then using that same reconstructed flight path as the experimental benchmark is premature, before asking, “How difficult was it for pilots to reproduce this manoeuvre?”One must first establish with reasonable confidence“Is this the manoeuvre that actually occurred?”Otherwise the study risks measuring the difficulty of reproducing a model of the event, rather than the event itself.The study’s own conclusions go beyond its experimental evidenceThe report concludes that the pattern of evidence fits better with “automated” control than “manual” control and describes automated precise targeting as a coherent explanation for the reported flight paths.But “coherent explanation” is not the same as “demonstrated explanation.”This is the central evidential problem, that the study has demonstrated, at most, that selected pilots found selected reconstructed manoeuvres difficult in a particular simulator under particular instructions. It has not demonstrated that the original aircraft had the same handling characteristics as the simulator, the reconstructed trajectories are exact, the pilots in 2001 had no relevant transferable skills, the original aircraft were controlled under identical conditions, the manoeuvres were actually as difficult in the real aircraft, manual control could not have produced the observed trajectories, or an automated system actually existed and was capable of producing those trajectories.The last point I want to highlight, which is particularly importantEvidence against one hypothesis is not automatically evidence proving a competing hypothesis. If manual control has not been conclusively established, the appropriate scientific conclusion is not automatically “therefore automated control.” There must be positive evidence for the alternative mechanism.The study’s treatment of automation is therefore asymmetricalThe experiment investigates manual control empirically but does not perform an equivalent experimental demonstration of the proposed automated-control mechanism. The researchers do not demonstrate a specific identified automated system reproducing the complete trajectories under the relevant conditions. Instead, automation functions partly as an explanatory alternative. This produces an asymmetry. Manual hypothesis, tested through pilot simulation.Automated hypothesis, primarily inferred from the difficulty of manual reproduction and the apparent precision of the reported flight paths.That is not an apples-to-apples comparison. A genuinely discriminating experiment would test both hypotheses under equivalent conditions. For example, researchers could identify a specific automated guidance system, model it accurately, establish its availability at the relevant time, establish its operational capabilities, and demonstrate that it could reproduce the observed flight paths.Without that, automation remains an explanation rather than an experimentally demonstrated cause.The study has confirmation biasThe study’s framing is also important, its executive summary states that the findings support “automated control” and subsequently connects those findings to a much broader conclusion concerning 9/11 as a “false flag event.”That does not automatically invalidate the experiment. Researchers are permitted to have hypotheses, however, when a study begins with a contested hypothesis and ends by connecting its experimental results to a much broader predetermined interpretation of historical events, the burden of methodological neutrality becomes especially important. The experiment should therefore be evaluated independently of the broader narrative.The question should be; What does the simulator experiment actually demonstrate Rather than; Does the simulator experiment support the wider theory? Those are very different questions.The study contains a particularly important internal contradictionThe authors acknowledge that their simulator may have made the manoeuvres more difficult than they would have been in a real 767 or 757. They further acknowledge that different simulated aircraft and different instructions could produce higher completion rates, yet they subsequently describe the observed results as providing substantive evidence regarding the difficulty of the manoeuvres and use that difficulty as an important component of the argument against manual control. This creates a tension between the study’s limitations and its conclusions. If the principal experimental variable is potentially biased in the direction of making the manoeuvre harder, then the resulting failure rate should be treated as an upper-bound indication of difficulty, not necessarily as an accurate representation of real-world feasibility.The correct scientific response to such a limitation would be to repeat the experiment using aircraft-specific simulators or validated aerodynamic models before drawing strong conclusions.What the study can legitimately conclude?A more defensible conclusion would be considerably narrower. The study provides evidence that: Under the particular experimental conditions employed, including use of a Boeing 737simulator, the selected pilots generally found the reconstructed indirect manoeuvres more difficult than simpler direct approaches. That is a reasonable experimental observation. The study may also legitimately argue that the reported trajectories appear more complicated than alternative trajectories that would have accomplished the same broad objective, but, that does not establish that the original manoeuvres were impossible, improbable, or necessarily automated. Neither does it establish that the people controlling the original aircraft could not have performed them, and it certainly does not establish the existence or use of a particular automated guidance system.My ConclusionThe fundamental weakness of Manual or Automated, is therefore not simply that it used a Boeing 737simulator instead of a Boeing 767 or 757. That is only the most obvious manifestation of a much broader methodological problem. The study attempts to infer the probability of a historical event from a small number of simulator experiments in which pilots were asked to reproduce reconstructed flight paths using an aircraft model that the researchers themselves acknowledge may have had different manoeuvrability characteristics from the aircraft actually under investigation.At the same time, the experiment contains uncertainties concerning starting position, timing, instructions, graphics, simulator familiarity and reconstruction of the original flight paths. The researchers explicitly acknowledge many of these limitations.The experiment therefore does not isolate the variable it needs to isolate, it measures the performance of selected pilots in a particular simulator, under particular instructions, attempting to reproduce a reconstructed trajectory. That is not equivalent to measuring whether a Boeing 767 or 757 could have flown the original manoeuvre manually.Most importantly, even if the study had demonstrated that the manoeuvres were extraordinarily difficult, that would still not prove automation. At most, it would establish that manual execution was difficult under the tested circumstances. A separate body of positive evidence would be required to establish that an automated or remote-control system actually performed the manoeuvres.The study’s own acknowledgements make this distinction particularly important. It concedes that different simulators and instructions might produce different completion rates, that further investigation of the underlying flight-path evidence is warranted, and that important questions concerning FDRs, radar tracking and ACARS data remain open.Consequently, the strongest criticism is not that the study has “proved the opposite” of its conclusion, it is that the experimental design does not provide a sufficiently valid or discriminating basis for the magnitude of the conclusion claimed. In scientific terms, the study has a problem of external validity, construct validity, experimental control, small-sample inference and causal identification. The Boeing 737/767 mismatch is particularly significant because it potentially biases the central dependent variable—the difficulty of performing the manoeuvre. Until that problem is resolved through aircraft-specific simulation or a rigorously validated performance model, the reported failure rates cannot reasonably be treated as reliable measurements of the difficulty that a Boeing 767 or 757 pilot would have experienced.The most appropriate and truthful conclusion is therefore one of substantial uncertainty, not proof of “automated control”.The study may be useful as an exploratory simulation exercise. It is much less persuasive as a controlled scientific experiment capable of distinguishing manual from automated flight control.Thanks for reading and caring!
Archived URL: https://911planesresearch.substack.com/p/why-the-manual-or-automated-study
�� CONTENT HASHES:
SHA-256: 9c297f45da666c7da01957da0a9d9f1a440df7b4415e3842e0e4abf6c4bdcfad
BLAKE2b: c5025a115d0507b5cb791554809801da7d8ce84e8f45c07e57a0924db505258b
MD5: fe44706e12b5fd390b06ce76299e39d1
�� TITLE HASHES:
SHA-256: 2286ee13bcd752f60a9bcef1e56529c87d27020f805629a044115d6cf4ae1ff0
BLAKE2b: ac0cafa49e75a10becd3e28362ab5af2c557b788ae52f3692947c9cde4182cbc
MD5: 0587d0605ff838c5ff5944fdf74e5ff7
�� INTEGRITY HASHES:
SHA-256: ec56e1d81371204685fbaae0b77070635c7f6573253da8b255d0a26e953a7731
BLAKE2b: c0f36465dc757f9653d5122c9150a14f3e708ef176859df51a62640c6a198054
MD5: 216975adc8874e492b517a2fef3616a0
Archived with ArcHive - Client-side cryptographic archival system