Source: https://911planesresearch.substack.com/p/a-forensic-page-by-page-critique
Following-on from my previous article, Why the “Manual or Automated?” Study Is Fundamentally Flawed, I decided to deploy ChatGPT to conduct a full forensic, page-by-page critique of the 51 page study, produced by Dr. Piers Robinson, in May 2026.I would advise readers to read the study for themselves, so they can fully understand the problems which I pointed out in my previous article. The construction of the study itself was set-up to give a false, flawed binary choice, and conclusion. This article covers a ChatGPT unbiased request to analyse the study, and report its results, which I have published here. A Forensic, Page-by-Page Critique of the Manual or Automated? A Flight Simulation Study and Analysis of Reported Aircraft Maneuvers on September 11, 2001 The study’s central methodological problem can be stated very simply:It attempts to determine whether Boeing 767 and 757 aircraft could have been manually flown along particular reconstructed trajectories, but the principal simulator experiment was conducted using a Boeing 737 simulator whose handling characteristics the authors themselves acknowledge may differ materially from those of the aircraft being investigated.That problem is compounded by:Extremely small sample sizesArtificial starting conditionsResearcher-controlled timing cuesChanging instructions between experimental stagesDeliberate omission of parts of the manoeuvresSimulator familiarisation effectsReconstructed rather than directly measured flight pathsAmbiguous coding of some resultsReliance on first-attempt performanceLack of an equivalent experimental demonstration of the proposed automated mechanismA substantial inferential leap from “difficult manually” to “therefore automated.”Most importantly, the study repeatedly moves from what the experiment actually measured to what the authors believe probably happened historically. Those are not the same thing.Pages 1–2 — Authorship, credentials and framingPages 1–2 identify Piers Robinson as the author and research director of IC911, with a PPL/IMC rating and approximately 200 hours of flying experience. Seven pilots participated in the simulations, alongside contributors and reviewers.Methodological issueThe author’s 200 hours of flying experience is not itself a problem. Nor is institutional affiliation evidence that the study is incorrect.But the report subsequently makes highly specialised claims about aircraft handling, flight-path reconstruction, probability and automated guidance.The study therefore requires particularly strong methodological safeguards.What is missing from these opening pages is an independent statement concerning:Independent peer reviewSimulator validation against the actual aircraftIndependent reconstruction of the flight pathsPre-registration of hypothesesStatistical analysis planOr independent replicationThis becomes important later because the report’s conclusions are considerably broader than its experimental design.Page 3 — The first major logical problemThe executive summary states that direct routes to the targets were readily accomplishable, while the reported routes involved “unnecessarily demanding” manoeuvres.This is already an important methodological distinction. The experiment demonstrates that alternative routes were easier. It does not demonstrate that the reported routes were impossible or implausible. An aircraft can follow a more difficult route for many reasons.The study needs to establish:Why should the existence of an easier route be evidence against manual flight?That proposition is never experimentally established.A pilot does not necessarily choose the geometrically simplest possible route merely because it is easier.The study effectively introduces an unstated behavioural premise:A manually controlled pilot would necessarily choose the easiest available route.That is an assumption, not a measured law of aviation behaviour.Page 4 — The conclusion appears before the evidencePage 4 goes substantially further.It argues that the reported flight paths show deliberate precision and that the simulations support automated control.The crucial problem is the phrase:“Taken together, the flight simulations and analysis of available flight path data support the hypothesis of automated control.”The experiment does not actually test an automated-control system.It tests human pilots flying a 737 simulator.Thus the experiment provides data concerning the difficulty of a manually performed task. It does not provide corresponding experimental data demonstrating the operation of an automated system. That creates an asymmetrical comparison:Manual: Experimentally tested.Automated: Inferred.That is not a controlled comparison between two competing hypotheses.Pages 4–5 — The study crosses from aviation analysis into a much larger historical conclusionThe report then states that its findings strengthen the thesis that 9/11 was a false-flag event rather than a terrorist attack.This is an enormous evidential leap.Even if the aviation experiment were completely sound, demonstrating that a particular manoeuvre was difficult for a sample of simulator pilots would not establish:Who controlled the aircraftWhether remote control existedWhether an automated guidance system existedWhether such a system was installedWhether it was activatedWho operated itOr why?The experiment therefore cannot bear the evidential weight assigned to it in the broader conclusion.Pages 5–6 — Aircraft mismatch becomes centralThe study identifies UA175 and AA11 as Boeing 767s and AA77 as a Boeing 757.Then, on page 8, it reveals the experimental aircraft:“a full motion (Category D equivalent) Boeing 737 flight simulator was secured for use.”This is arguably the single most important methodological weakness.A Boeing 737 is not a Boeing 767.Nor is either aircraft a Boeing 757.The relevant question is therefore not whether both are commercial Boeing jets.The relevant question is whether the 737 simulator accurately reproduces the specific aircraft-performance characteristics that determine the manoeuvre being studied.Those include:roll responsepitch responsecontrol sensitivityinertiaaerodynamic responsehigh-speed handlingturn performanceenergy retentionthrust responseand control-system behaviourThe report does not provide a quantitative validation demonstrating equivalence in these characteristics.Page 8 — The “validation” of the 737 is inadequateThe authors attempt to deal with the aircraft mismatch.Two highly experienced pilots, each with more than 10,000 hours, evaluated the simulator. They did not actually fly the 9/11 manoeuvres. They assessed handling at high and low speeds and concluded that it was a reasonable approximation of a 767 or 757.This is not a rigorous validation.A qualitative opinion that:“this feels reasonably similar”is not equivalent to establishing:“the simulator reproduces the relevant 767/757 aerodynamic and control characteristics within a defined error margin.”That distinction is critical.The study needed to establish fitness for purpose, not merely general plausibility.And later in the report, the researchers themselves admit that the 737 may have been less manoeuvrable than the actual 767 or 757.That admission substantially weakens the earlier validation.Page 9 — The first-attempt premiseThe study decides that first attempts are particularly important because the aircraft supposedly hit the targets on their first attempts.This is a major assumption.The authors argue that prior simulator practice would not have provided knowledge of how the real aircraft handled.But the relevant question is not simply:Did the pilot practise this exact manoeuvre?It is:What transferable aviation skills did the original pilots possess?A pilot does not learn every manoeuvre from scratch.Pilots transfer skills between aircraft and situations.The experiment’s own results demonstrate the importance of familiarity with the particular simulator.That creates a confounding variable:failure may reflect unfamiliarity with the simulator rather than inability to perform the manoeuvre.Page 9 — Another problem: 15 minutes of familiarisationThe study states that the pilots received at least 15 minutes to practise handling the simulator before commencing runs.This is problematic because the later pilot testimony explicitly identifies simulator unfamiliarity as a major factor.Pilot 6 said that the difficulty involved his unfamiliarity with the 737 simulator’s responses and that he believed he could adapt more sharply in an aircraft he normally flew.Pilot 3 similarly described unfamiliarity with the control inputs and differences between aircraft control systems.This is extraordinarily important.The study is effectively using failure caused partly by unfamiliarity with the experimental aircraft as evidence concerning the ability to fly a different aircraft.That is precisely the confound that should have been eliminated.Pages 10–11 — Direct routes are not a control for manual capabilityThe direct-route experiments show that pilots could easily hit the targets.That is useful.But the study interprets this as evidence that a manually controlled aircraft would have taken the direct route.That conclusion does not follow.The experiment demonstrates:The direct route was easier.It does not demonstrate:A manually controlled aircraft would necessarily take the easiest route.Those are completely different propositions.This is one of the study’s recurring problems: an empirical observation is converted into a behavioural assumption.Pages 12–13 — The experiment is not actually a complete replicationThe South Tower experiment is particularly problematic.Pilots were given a command approximately ten seconds before impact telling them when to initiate the turn.This means the experimental pilot was given information that the original pilot supposedly did not have.The researcher knew:the targetthe desired trajectorythe timingthe intended manoeuvreand when the turn was supposed to beginThe original pilot, by contrast, was not operating under experimental instructions.More importantly, the study explicitly says:“we did not give pilots instructions to initiate the last-second pull-out and instead focused on the execution of the final turn only.”This is a fundamental limitation.The experiment did not reproduce the complete manoeuvre.Yet the conclusions discuss the difficulty of the complete last-second manoeuvre.That is a classic construct-validity problem.Page 14 — A potentially serious heading errorThe study says it identified an error in the NTSB’s conversion of true headings to magnetic headings and corrected it.This is potentially important.If the researchers are correcting the source data used to construct the experimental trajectory, then the study must establish:precisely what the original error washow it arosehow the correction was calculatedwhether all subsequent coordinates were recalculatedwhether the NTSB’s underlying radar data were affectedand whether the corrected path is independently validatedThe report does not provide sufficient detail in the study itself to establish the uncertainty introduced by this correction.This matters because small heading differences at high speed can produce substantial positional differences over seconds.Pages 15–17 — The Pentagon experiment introduces additional artificial constraintsThe Pentagon experiment is even more complicated.Pilots were instructed to reproduce the 330-degree orbit and low-level approach.The instructions were then modified.At one point pilots were told to attempt to clip light poles; later the instructions were changed so that the poles became visual guides.This creates a major experimental-design problem:The experimental conditions were not constant.If instructions change between participants or stages, then performance cannot simply be treated as though all participants were operating under identical conditions.The experiment is no longer testing merely:“Can pilots perform this manoeuvre?”It is testing:“Can pilots perform this manoeuvre under this particular set of instructions?”Those are not equivalent questions.Pages 18–22 — The direct-route results actually establish something usefulThe direct-route experiment is arguably the strongest part of the study.All four lower-experience pilots successfully aligned with the South Tower, and three of four successfully struck the Pentagon directly.This demonstrates something valuable:Hitting a large stationary target with a large aircraft in a simulator is not necessarily difficult merely because the aircraft is travelling fast.But it does not establish the subsequent proposition:Therefore the original pilots would necessarily have chosen the simplest trajectory.That remains an inference.Pages 23–24 — The most damaging experimental admissionPage 23 contains one of the most important passages in the entire study.The authors explicitly acknowledge:uncertainty about the correct offsetuncertainty about when to instruct the turnpossible lower manoeuvrability of the 737graphics limitationsproblems with the instructionsand the possibility that other simulators could produce higher completion ratesThis is not a minor limitations paragraph.These are precisely the variables that determine whether the experiment is measuring what it claims to measure.The report nevertheless says the results can still support substantive conclusions.That is where the methodological argument becomes weak.If the apparatus itself may systematically increase the difficulty of the manoeuvre, then the failure rate cannot be treated as an unbiased estimate of real-world difficulty.Pages 23–26 — The results reveal the simulator-learning problemThe study reports zero first-attempt completions for the South Tower indirect manoeuvre.But the following pages explain why.The pilots had difficulty determining:appropriate roll angleappropriate back pressurecorrect initiation timingand how the simulator respondedCritically, the pilots themselves describe unfamiliarity with the simulator.Pilot 2 said:“The roll rate is quite low.”and explicitly referred to getting used to how the simulator performed.Pilot 3 described unfamiliar control fidelity and stated that he was unfamiliar with the control inputs despite having thousands of hours in aircraft.Pilot 6 was even more explicit: he said his difficulty resulted from unfamiliarity with the 737 simulator’s response and that he believed he could adapt the turn more sharply in an aircraft he normally flew.This is perhaps the strongest evidence against treating the first-attempt failures as direct evidence of real-world impossibility.The pilots themselves identify simulator unfamiliarity as a cause of failure.Page 27 — The study turns pilot preference into evidenceThe pilots repeatedly say that they would personally have flown a more straightforward route.This is interesting qualitative evidence.But it is not proof of what another pilot would have done.Nor does it establish that a pilot choosing another route is evidence of automation.This is essentially an appeal to expert preference:“I would not fly it that way.”That is not the same as:“A competent pilot could not fly it that way.”The distinction is crucial.Page 28 — The critical inferential leapThe study concludes that the required precision was inconsistent with manual control by the alleged hijackers.But the experiment has not established the probability of successful manual execution.It has established a small experimental failure rate under specific conditions.Those are not equivalent.There is no statistical model that converts:0/5 first attemptsinto:probability that the original pilot could not perform the manoeuvre.Nor is there a Bayesian analysis incorporating:prior pilot experienceaircraft differencesroute uncertaintyactual visual conditionspossible navigation assistanceand uncertainty in the reconstructed trajectoryThe phrase “inconsistent with manual control” therefore goes beyond the experimental data.Pages 28–30 — Pentagon completion rate: 28%The study reports 25 attempts at the Pentagon low-level manoeuvre, of which seven succeeded, giving a 28% completion rate. None succeeded on the first attempt.At first glance, this appears powerful.But the 25 runs are not a clean statistical sample.They include:different pilotsdifferent attemptsrepeated attempts by the same pilotsdifferent starting pointsevolving pilot experiencemodified instructionsand different experimental conditionsTherefore the 25 runs cannot simply be treated as 25 independent observations.This is a major statistical issue.If Pilot 1 performs five runs, those five observations are not equivalent to five independent pilots.The study does not provide an appropriate statistical treatment of this repeated-measures structure.Consequently, quoting 28% completion creates an appearance of statistical precision that the experimental design does not support.Pages 29–30 — Visual acquisition is confounded with flight skillThe study says pilots had difficulty reacquiring the Pentagon visually after the orbit.But this is not necessarily evidence of the difficulty of manually controlling the aircraft.It is partly a test of:How easily can a simulator pilot visually reacquire a target after being deliberately instructed to fly away from it?That depends upon:simulator graphicsfield of viewrenderingcockpit geometryvisual fidelityatmospheric representationand pilot familiarity with the simulated environmentThe report itself acknowledges that graphics quality may have complicated the Pentagon runs.Therefore visual reacquisition failures cannot simply be interpreted as evidence that the original manoeuvre was beyond manual flying capability.Pages 30–33 — The study effectively admits that the manoeuvre was abnormalThe pilots repeatedly say the manoeuvre was unnatural, counter-intuitive and difficult.That is not surprising.They were explicitly instructed to reproduce a highly unusual trajectory.Pilot 1 even described it as something outside normal piloting behaviour.But again:Unusual does not mean impossible.The study repeatedly conflates:unusualdifficultcounter-intuitiveunlikelyand impossibleThese are different propositions.A manoeuvre can be:possible + difficult + unusual + successfully executed.That combination is entirely compatible with manual flight.Pages 34–36 — Reconstruction is treated as established factThe study reconstructs UA175 and AA11 trajectories from NTSB headings, radar and film analysis.But the reconstruction itself is a model.The researchers are not measuring the original control inputs.They are reconstructing a path from available observations.That means the experiment actually contains two models:Model 1: reconstruction of the historical flight path.Model 2: simulation of a pilot attempting to reproduce Model 1.Any uncertainty in Model 1 propagates into Model 2.Yet the study’s conclusions frequently speak as though the reconstructed path were an exact record of the original aircraft’s control trajectory.That is an important epistemological problem.Page 37 — “Perfect perpendicularity” is treated as stronger evidence than demonstratedThe study argues that AA11 achieved an almost perfectly perpendicular impact trajectory and that this is strongly suggestive of automated control.There are several problems.First, the reported trajectory itself has uncertainty.Second, an impact angle close to perpendicular does not inherently identify the control mechanism.A manually flown aircraft can produce a straight final trajectory.The relevant question is not:“Could automation produce this?”Obviously it could, assuming an appropriate system.The relevant question is:“What is the likelihood of obtaining this trajectory manually versus automatically, given all known information?”The study does not actually calculate those competing probabilities.It instead asserts that the symmetry is “strongly suggestive.”That is evidentially weaker than the language used elsewhere in the report.Pages 38–40 — The “symmetry” argument contains an important assumptionThe report argues that because both aircraft eventually approached the towers at approximately perpendicular angles, the probability of this occurring under manual control was extremely low.But no actual probability calculation is presented.This is therefore a qualitative probability assertion.To claim that the probability is “extremely low,” the researchers would need a model defining:the distribution of possible approach angles under manual controlpilot behaviourtarget geometryavailable visual cuesnavigation informationaircraft performanceand the probability of correcting toward a targetNone of this is quantified.Therefore:“extremely low probability”is asserted rather than demonstrated.Page 40 — The terminal-guidance hypothesis is introduced without establishing the systemPage 40 is where the argument shifts decisively from observation to speculation.The authors propose that a terminal guidance system may have “kicked in” during the final seconds.But the study does not establish:the existence of such a systemits architectureits availability in 2001whether it could control a 767/757whether it could identify the targetwhether it could generate the required bank/pitch inputswhether it could operate at the reported speedsor whether such a system was installed on the aircraftThus the proposed mechanism is hypothetical.It cannot serve as the demonstrated explanation for the simulator results.Pages 41–43 — The “manual versus automated” table is not actually a testThe study produces a table categorising findings as either consistent with manual or automated control.This creates the appearance of a formal hypothesis test.But the categories are qualitative.There is no:likelihood ratioconfidence intervalprobabilitystatistical testBayesian posteriorerror estimateor quantitative comparison between competing hypothesesThe table therefore represents the authors’ interpretation rather than a statistical test.Pages 43–45 — The alternative hypothesis is unfairly constructedThe study presents the manual explanation as requiring multiple unlikely propositions.But notice how the alternative is constructed.It assumes:the pilots deliberately chose convoluted routesthe routes were mistakesthey placed themselves in highly difficult positionsthey then successfully recoveredand they independently achieved similar impact geometryThe problem is that this is not the only manual-control hypothesis.There are other possibilities.For example:The pilot deliberately flew a particular route for reasons unrelated to minimising the geometric distance to the target and then successfully controlled the aircraft during the final approach.The study does not adequately model the full range of possible manual-control behaviours.It constructs a relatively unattractive version of the manual hypothesis and compares it with a highly purposeful automated hypothesis.That introduces hypothesis asymmetry.Pages 45–46 — “Planning would have made it simpler” is not demonstratedThe study argues that if the routes had been planned and practised, the pilots would have chosen easier trajectories.Again, this is a behavioural assumption.Planning does not necessarily produce the mathematically simplest route.A route can be selected for many reasons:navigationtarget orientationconcealmenttiminggeographic constraintsvisual acquisitionor simply pilot preferenceThe study does not establish that the alleged pilots were optimising route simplicity.Pages 46–47 — The conclusion exceeds the evidenceThe conclusion says the pattern of evidence fits better with automation than manual control.“Fits better” is a legitimate scientific formulation if the competing models have actually been quantitatively compared.Here they have not.Instead, the conclusion rests upon a chain:737 simulator failures, reported manoeuvres were difficult the manoeuvres were counter-intuitive pilots would probably have chosen easier routessimilar final geometries are suspicious and therefore automated guidance provides a better explanation.Each, introduces an assumption.The study does not independently validate every link.Pages 47–48 — A particularly revealing contradictionThe final section acknowledges that further research is needed into:the missing FDRsAA77’s FDRFDR interpretationUA93earlier flight pathsAA77’s radar trackingACARSpossible guidance systemsand other evidenceThis is extremely important.The study therefore effectively says:The evidence necessary to resolve several of the central questions is still unavailable or requires further investigation.Yet the earlier sections present the automated-control conclusion with considerably greater confidence.That is an internal tension.If further investigation is required to determine whether the FDR data can distinguish manual from automated control, then the study should be cautious about claiming that its own simulator results establish automated control.Page 49 — Source selectionThe reference list includes a mixture of:NTSBNIST9/11 Commission materialacademic publicationsdocumentary materialadvocacy organisationsYouTube materialand sources advancing competing 9/11 interpretationsThe problem is not that controversial sources are cited. The problem is that the study does not consistently distinguish between:Primary evidence and interpretations of primary evidence.For a study whose central purpose is to determine aircraft performance, the hierarchy should strongly favour:original radar dataoriginal FDR dataaircraft performance datavalidated simulator modelsindependent aerodynamic analysispeer-reviewed aviation researchInterpretive material should occupy a secondary position.Page 51 — The appendix creates another reproducibility problemThe study says that videos of the simulator runs are available through hyperlinks.That is useful, but video observation alone does not provide the underlying experimental dataset.A reproducible experiment should ideally provide:exact starting coordinatesaltitudeairspeedaircraft weightcentre of gravitywindtemperature;atmospheric modelengine settingscontrol inputsbank anglespitchvertical accelerationheadingsimulator model/versionsoftware versionsampling rateand precise success/failure criteriaWithout those data, independent researchers cannot fully reproduce the experiment.The five most serious methodological flawsAfter reviewing the entire report, I would rank the problems as follows.1. Aircraft-model validityThis is the most fundamental.The study investigates 767/757 manoeuvres using a 737 simulator.The authors themselves acknowledge that the 737 may have been less manoeuvrable than the aircraft actually under investigation.That means the independent experimental apparatus may systematically bias the result.2. The experiment does not test the complete manoeuvresThe South Tower experiment explicitly omits the last-second pull-out.Therefore the experiment cannot legitimately claim to reproduce the complete manoeuvre.3. First-attempt failure is confounded by simulator unfamiliarityThe pilots themselves repeatedly identify unfamiliarity with the 737 simulator as a source of difficulty.That directly undermines the use of first-attempt failures as evidence about real-world first-attempt performance in a different aircraft.4. The study does not establish the probability of manual successThe study reports completion rates but does not establish a statistically valid probability that the original pilots could or could not have performed the manoeuvres.The 28% Pentagon figure, for example, comes from repeated runs involving the same pilots under varying conditions.Those are not 25 independent trials.5. Failure of manual reproduction does not prove automationThis is the fundamental logical error.The alternatives include:manual controlmanual control plus navigation assistancemanual control by a more capable pilotprior knowledge or practicedifferences between aircrafterrors in the reconstructed trajectoryor simply successful manual execution of a difficult manoeuvreThe study does not experimentally eliminate these alternatives. Nor does it demonstrate an actual automated system reproducing the historical flight paths.The fundamental evidential chain is therefore incompleteThe study essentially needs to establish this:A. The reconstructed flight path is correct.B. The 737 simulator accurately represents the relevant 767/757 characteristics.C. The experimental instructions faithfully reproduce the original circumstances.D. The pilot sample represents the relevant population.E. First-attempt simulator performance accurately predicts real-world first-attempt performance.F. The observed failure rate establishes that manual execution was highly improbable.G. No alternative manual explanation adequately accounts for the trajectory.H. A specific automated system can reproduce the trajectory.I. That system existed and was available on the aircraft.J. Therefore automation was the most probable cause.The study does not establish this entire chain.At several points it effectively jumps from A–F to J.The strongest criticism of the studyThe most defensible criticism is therefore not:“The study proves the official account.”Nor should the criticism simply be:“The study is wrong because it used a 737.”That would be too simplistic.The much stronger methodological argument is:The study’s experimental apparatus, experimental conditions and inferential framework are insufficiently validated to support the strength of the conclusions drawn from them.The 737/767/757 mismatch is especially important because it directly affects the quantity being measured: manoeuvrability and pilot control difficulty.The researchers acknowledge this limitation themselves.They also acknowledge uncertainty concerning the starting offset, timing, graphics and instructions, and concede that different simulators and instructions could produce higher completion rates.Once those limitations are recognised, the simulator results become exploratory evidence, not decisive evidence, and that distinction is critical.What the study actually demonstratesA cautious interpretation would be:The selected pilots generally found the reconstructed indirect manoeuvres difficult to reproduce under the particular experimental conditions used.What the study claimsIt moves considerably further:The reported manoeuvres were unlikely to have been manually flown and are better explained by automated guidance.Those propositions are not equivalent.The first is supported by the experiment.The second requires additional evidence that the experiment itself does not provide.That is where I believe the study’s fundamental methodological weakness lies.Thanks for reading and caring!
Archived URL: https://911planesresearch.substack.com/p/a-forensic-page-by-page-critique
�� CONTENT HASHES:
SHA-256: 7cb23219e36e61735588b37d8db0693febc0e3194b85e24324f7c707a1c32736
BLAKE2b: ffc1a1998ac6bcd2c93f308617a47bce3030fbc6610cbb7444072836ddef4ed2
MD5: a1920c47d698d4387319aae115710fdf
�� TITLE HASHES:
SHA-256: 0657d3c0c7e8afba6ca55f2700ab0595adf870feaf5332f24613d9fcedc844db
BLAKE2b: c534a97d01073e66174f038a3cc83680a0a9c0e29eb093fe041da2c48497a634
MD5: 9c5bcce248242c32f3e8203a38147821
�� INTEGRITY HASHES:
SHA-256: c0c6ee7339b9003b3ba6443c36cde7ec37c70742cf2cac080681a05aa4b4175d
BLAKE2b: 4625d70eb37b62c4db89ab23e91775e799380c5ac689450c623f3c02cc2a3cfc
MD5: 42a7430593e275fd7b9a95babb18f8e6
Archived with ArcHive - Client-side cryptographic archival system