Air traffic control and medical devices

“It’s Déjà Vu All Over Again”

A recurring UK traveler nightmare

United Kingdom (UK) flights disrupted … again.

On September 21st, disruptions again are affecting flights in/to/from the UK. As in the major September 8th outage, the cause is a “technical issue” in the National Air Traffic Services (NATS) systems that provide UK air traffic control.

A 1 millisecond event caused the prior September 8th NATS outage. Recovery of normal NATS operations took around 6 hours. The effects caused over 2,000 flight cancellations and more than two days for air travel recovery.

The September 8th event investigation (https://www.nats.aero/news/nats-publishes-preliminary-report-on-technical-incident-of-8-september/) currently pins the cause on a failure for a paused operation to resume correctly. The failure to resume correctly created corrupted output which then caused system-wide issues. Pausing lower priority or other operations briefly is not unusual. Not resuming normally, however, should be unusual. The defect, called by some a “rare sequence” problem, had been in the software for some time.

Why is this interesting from a medical device software perspective?

Medical device disruptions can be caused by similar issues.

If there is an ability to pause and resume medical device software functions, then any pause-resume sequences will likely need both normal mode and abnormal mode testing. Test cases would also have diversity scaled to the risk posed by a resumption error.

An example: an infusion pump has an alarm feature for notifying a health care professional (HCP), e.g., reservoir low that can endanger the patient. The alarm software functionality also allows for the HCP to pause the alarm while resolving the critical event. What if the alarm no longer works after pause? What if the alarm, while paused, preempts subsequent alarms?

Two common ways to attack this type of disruption

There are lots of ways to reduce or prevent software defects. Here are two, one an architectural approach and the other a software detail design approach:

  • Architectural diagrams can increase visibility and understanding of the design. Sequence, state, or event-driven architectural diagrams can make identification of key faults much easier.
  • Input data checking. The perhaps far more concerning problem in the September 8th NATS outage was in software “downstream” from the defect. That software apparently did not have adequate checks of input data to detect the corruption. Nor did the downstream functions have an ability to continue operations if corrupted data was received. Implementing input checks and providing exception processing capabilities is not “cutting edge” engineering. Considering those two aspects should be part of detailed design for any software function.

Don’t wait to see if forgotten or ignored software engineering fundamentals affect your medical device.

About the author

Succeeding despite relentless change is the goal of 21st century organizations. Mike helps achieve those successes by working with leaders of start-ups to Fortune 20 companies and national governments. His aim is to help them re-imagine and create adaptive, innovative enterprises that increase profitability and value across the quadruple bottom line: customers, employees, owners/shareholders, and communities.

Upcoming Public Courses

No public courses are planned at this time.

If you would like to discuss a private, onsite course, please fill out the form below.

Or just email training@softwarecpr.com for more info.

Corporate Office

15148 Springview St.
Tampa, FL 33624
USA
+1-781-721-2921
Partners located in the US (CA, FL, MA, MN, TX) and Canada.