1. Marietta Incident (Kennestone, June 1985)
Woman severely injured (paralysis in some areas)
Operators and manufacturers did not believe the machine wasn’t working (wasn’t even a possibility) Lack of documentation (machine reports or otherwise) Accident was not acknowledged or announced by the manufacturer AECL.
2. Ontario, July 1985
Seven weeks after Kennestone incident
Incorrect error message displaying NO DOSE and TREATMENT PAUSE The operator then tried the machine again 4 times with the same message. Stated to not be an unusual case by an operator. 6 more patients treated that day. First patient hospitalized, then machine taken out of service. Service engineer sent and users were informed of an issue (not that an injury had been caused, however).
3. Manufacturer and Government Response
AECL could not reproduce the malfunction, but after assumptions were made about the conditions, they stated they had improved it by “5 orders of magnitude,” but in their actual government report, they stated they could not exactly state what the cause of the accident was.
There is reasonable suspicion that it was a software error and not a microswitch failure. There was a voluntary recall by AECL, but later the FDA audited AECL’s modifications and users were told to resume normal procedures.
Canadian government report indicated that AECL must implement hardware and software changes, including an automatic halt on dose-rate malfunction. AECL only implemented the microswitch changes, allowing up to 3 treatment retries, and also did not comply with the suggestion to use an additional independent system.
Issues so far
Manufacturer lack of accountability Lack of clear communication Lack of rigorous debugging and further testing before redeployment Uninformed operators (knowing when to step, enforcement, etc)
Yakima Valley, Washington
Machine malfunction still not acknowledged until after later accidents. A woman developed a visible symptom but still continued to be treated.
AECL tech support supervisor said it could not have been a result of their machine, even stating that there has been no other accidents of similar damage.
The staff stated to not believe it was the fault of the machine either at the time due to the grossly exaggerative improvement claim by AECL.
East Texas Cancer Center
Easily adjustable settings with large consequences. Typing wrong value could deal lots of damage. Incorrect and unclear error message (wrong severity classification) - doesn’t detail properly, or severity of issues properly Staff not able to see/hear patient. Patient died.
User and Manufacturer Response
Therac-25 shut down. AECL could not reproduce error. AECL engineer said it was not possible for the machine to overdose a patient. AECL personnel said no other accidents had happened. Therac-25 use resumed after ETCC physicist determined it was working.
ETCC 2
Same Malfunction 54 error message. Same operator. Audio now hearable. Patient died.
User and Manufacturer Response
ETCC Physicist and Operator were able to reproduce the Malfunction 54 error, eventually at will.
AECL said they could not reproduce until ETCC informed them of the condition.
Yakima 2
after rigorous FDA training, use resumed
user error; software doesn’t check for user errors. user could still proceed with treatment again incorrect user interface display (wrong number of rads); misinfo to user
Investigators could not reproduce error AECL QA staff said that earlier changes that had yet to be implemented would have actually prevented the error (why wasn’t there a longer rollback??).
FDA to finally order shutdown of all Therac-25s
Overall Issues
A lacking government approval procedure Users not receiving much transparent information Lacking QA/test protocol Lack of urgency (rigorous testing) Lacking error messages (resulting in hardship reproducing errors) Lack of third party evaluations (unbiased views) Lack of documentation (testing or otherwise)
Inadequate follow-up on accident reports (safety-critical systems and follow-up procedures (almost non-existent during Kennestone)) Unrealistic risk assessments
Generalization and Learnings
- Overconfidence in Software (assuming code cannot / will not fail)
- Confusing reliability and safety (code can work many times, but when it doesn’t, is it still safe?)
- Lack of defensive design - no self-checking, error detection or handling
- Protect against root causes, rather than eliminating surface issues (there will always be another bug, so safeguard for those cases)
- Complacency (waiting for bugs to happen or harm before fixing)
- Importance of software documentations, quality, testing, and user interfaces
- Overconfidence in Software Reuse (reuse is not synonymous with safety)
- Safe versus Friendly UI