In 1972 I moved from the Advanced Research Laboratory, where CDC was developing the STAR-100 system, to the Scope Operating System group at CDC’s Arden Hills plant in Saint Paul. CDC wanted me to manage a group implementing part of their new operating system. Ray Nienberg, the General Manager, was willing to let me try my methods on one of the modules. They chose something that wasn’t really essential but that would be nice to have: a deadlock-preventing resource scheduler which would ensure that all jobs submitted for execution would run to completion without any resource deadlocks. The new operating system was to be released to users in 10 months.
I explained to the people in my group how we were going to implement our module. An analysis of precedent revealed that the problem had already been solved. Edsger Dijkstra called it the banker’s algorithm, and Wikipedia deals with it here. Not being one to reinvent the wheel, I simply implemented the banker’s algorithm for our operating environment resources. Decision tables were used to ensure that our implementation was robust.
Our group did not use one second of computer time during the first 6 months. We constructed and reviewed decision tables, refining them as we went. When we felt that all the rules (the meaningful combinations of alternatives of all the variables) had been dealt with we created an exhaustive testing mechanism to test our implementation. This suite of tests served as a regression test ensuring that when we fixed programming errors we didn’t break anything that had previously worked. The tests were designed to run until an expected outcome didn’t match a program outcome. If this happened we stopped the computer immediately and dumped out a complete map of the memory of the program. Each test situation was constructed mechanically from the decision tables. The tests were exhaustive and tested every possible situation. Remember, we were running on what was then the world’s fastest computer, the CDC 7600 (since we didn’t use any floating point arithmetic we must have been running even faster than its 36 MFLOP capability). When completed, the entire test, which included thousands of rules, took less than 1 second to execute.
We turned our module over to Integration and Evaluation on the day we got right through to the end of the testing sequence. We had completed our task within 8 months, and were one of the few groups ready to release their module by the release date. Integration and Evaluation declared our module ready for release after a couple of weeks of testing. I asked to see the test that Integration and Evaluation had used to test our module. They had used a dedicated super-computer for 300 minutes to test our module. On examination we found that they had tested some situations hundreds of millions of times while many situations had not been tested at all. We walked them through our tables and our testing procedure, and they got very excited. They realized the value of what we were doing and immediately wanted to use a similar methodology for all testing. They weren’t able to sell it to project management.
After reviewing the progress of the entire project the company management decided to slip the release date by 3 months. By now everyone in our group had a very good feel for our solution. We revisited our tables and decided how we could change things. After we had discussed the improvements for a few weeks we went ahead, altered the tables, revised the tests, and implemented the improvements.
The most costly aspect of maintaining a complex system is support. Fixing the bugs and releasing them to the users is very expensive. Monthly meetings were held to see how the new operating system was doing following its release. Every module would have an outstanding number of known bugs. On average there would be about 20 bugs per module. The support programmers would fix them at about the same rate, but the number of outstanding bugs didn’t appear to go down. New bugs were being created at about the same rate as old ones were fixed. Our module was the exception; a couple of bugs were reported in the first two months and then the module remained stable for the life of that version of the operating system (a number of years). My group was disbanded and all its members were reassigned to other projects. The maintenance of our module was assigned to another support group. There was virtually no cost attributable to maintaining our deadlock-preventing resource scheduler for the life of the operating system.
No comments:
Post a Comment