2012-03-14

Putting the methodology’s feet to the fire

In 1972 I moved from the Advanced Research Laboratory, where CDC was developing the STAR-100 system, to the Scope Operating System group at CDC’s Arden Hills plant in Saint Paul. CDC wanted me to manage a group implementing part of their new operating system. Ray Nienberg, the General Manager, was willing to let me try my methods on one of the modules. They chose something that wasn’t really essential but that would be nice to have: a deadlock-preventing resource scheduler which would ensure that all jobs submitted for execution would run to completion without any re­source deadlocks. The new operating system was to be released to users in 10 months.

I explained to the people in my group how we were going to implement our module. An analysis of precedent revealed that the problem had already been solved. Edsger Dijkstra called it the banker’s algorithm, and Wikipedia deals with it here. Not being one to reinvent the wheel, I simply imple­mented the banker’s algorithm for our operating environment resources. Decision tables were used to ensure that our implementation was robust.

Our group did not use one second of computer time during the first 6 months. We constructed and reviewed decision tables, refining them as we went. When we felt that all the rules (the meaningful combinations of alternatives of all the variables) had been dealt with we created an exhaustive test­ing mechanism to test our implementation. This suite of tests served as a regression test ensuring that when we fixed programming errors we didn’t break anything that had previously worked. The tests were designed to run until an expected outcome didn’t match a program outcome. If this hap­pened we stopped the computer immediately and dumped out a complete map of the memory of the program. Each test situation was constructed mechanically from the decision tables. The tests were exhaustive and tested every possible situation. Remember, we were running on what was then the world’s fastest computer, the CDC 7600 (since we didn’t use any floating point arithmetic we must have been running even faster than its 36 MFLOP capability). When completed, the entire test, which included thousands of rules, took less than 1 second to execute.

We turned our module over to Integration and Evaluation on the day we got right through to the end of the testing sequence. We had completed our task within 8 months, and were one of the few groups ready to release their module by the release date. Integration and Evaluation declared our module ready for release after a couple of weeks of testing. I asked to see the test that Integration and Evaluation had used to test our module. They had used a dedicated super-computer for 300 minutes to test our module. On examination we found that they had tested some situations hun­dreds of millions of times while many situations had not been tested at all. We walked them through our tables and our testing procedure, and they got very excited. They realized the value of what we were doing and immediately wanted to use a similar methodology for all testing. They weren’t able to sell it to project management.

After reviewing the progress of the entire project the company management decided to slip the re­lease date by 3 months. By now everyone in our group had a very good feel for our solution. We re­visited our tables and decided how we could change things. After we had discussed the improve­ments for a few weeks we went ahead, altered the tables, revised the tests, and implemented the im­provements.

The most costly aspect of maintaining a complex system is support. Fixing the bugs and releasing them to the users is very expensive. Monthly meetings were held to see how the new operating sys­tem was doing following its release. Every module would have an outstanding number of known bugs. On average there would be about 20 bugs per module. The support programmers would fix them at about the same rate, but the number of outstanding bugs didn’t appear to go down. New bugs were being created at about the same rate as old ones were fixed. Our module was the exception; a couple of bugs were reported in the first two months and then the module remained stable for the life of that version of the operating system (a number of years). My group was disbanded and all its members were reassigned to other projects. The maintenance of our module was assigned to another support group. There was virtually no cost attributable to maintaining our deadlock-preventing resource scheduler for the life of the operating system.

No comments:

Post a Comment