An Alternative to Costly High Tech Contingency Plans Copyright 1992 CAUSE From _CAUES/EFFECT_ Volume 15, Number 3, Fall 1992. Permission to copy or disseminate all or part of this material is granted provided that the copies are not made or distributed for commercial advantage, the CAUSE copyright and its date appear,and notice is given that copying is by permission of CAUSE, the association for managing and using information resources in higher education. To disseminate otherwise, or to republish, requires written permission.For further information, contact CAUSE, 4840 Pearl East Circle, Suite 302E, Boulder, CO 80301, 303-449-4430, e-mail info@CAUSE.colorado.edu AN ALTERNATIVE TO COSTLY HIGH TECH CONTINGENCY PLANS by John E. Bates, Ivan G. Fuchs, and Alan Greenberg ************************************************************************ JOHN BATES is Director of Information Systems Resources at McGill University, responsible for administrative computing. Recently he also became responsible for McGill's business operations. He is a past president of the Canadian Information Processing Society, and is a member of its computer science accreditation council. E-mail: john@ums1.mcgill.ca. IVAN FUCHS is a principal of Valfox Consultants, Inc. He specializes in risk analysis, independent verification and validation, and software engineering disciplines. He spent twenty years in an academic environment, most of it as the director of the computer center of a major Canadian university. E-mail: ivanf@cc.mcgill.ca. ALAN GREENBERG is Director of Computing and Telecommunications at McGill University, responsible for all central computing as well as data and voice communications. He also sits on the management boards of CA*net, RISQ (the Quebec regional network), and NetNorth. E-mail: alan@vm1.mcgill.ca. ************************************************************************ ABSTRACT: Higher education administrators and MIS managers recognize the need for contingency and disaster recovery plans. Unfortunately, the perceived cost and effort involved, together with day-to-day operational pressures, tend to relegate such planning to a lower priority--if it happens at all. McGill University questioned the need for expensive, high tech solutions and instead focused on risk assessment and risk avoidance for central as well as distributed and departmental systems-- based on the theory that prevention is better than cure. The terms "risk assessment" and "contingency planning" generally conjure up the image of well publicized computer system disasters such as the AT&T network failure in New York, the effects of the 1989 San Francisco earthquake and of Hurricane Hugo in North Carolina, the Hinsdale, Illinois, switching station fire, the 1987 Montreal-based Steinberg grocery chain computer center fire, and the KGB-sponsored German hacker's story, documented in Cliff Stoll's The Cuckoo's Egg.[1] There are numerous other, less publicized, incidents that may have made the local news, but whose fame has died with the next sporting event. These include fires in campus administration buildings, stolen letterheads and signature stamps, stolen passwords, stolen computers, students who crossed the boundary from inquisitive to intrusive, major power failures, and altered student or payroll data. Every system manager has an armful of stories. In short, to paraphrase a colloquialism, disasters happen! Incited by the news media, occasionally harassed by auditors, in some cases spurred by directives from government agencies,[2] but mostly driven by their own feeling of responsibility, many MIS managers in large commercial organizations have implemented elaborate disaster recovery plans. Leading this group are insurance companies, banks, large retailers, and industrial concerns dealing with sensitive or confidential matter. Usually, the plans involve extensive front-end work and the contracting for hot or cold backup sites, duplicate processing units, alternate communication links, and other costly solutions. In enterprises where all business stops if the computer system is unavailable, such steps are clearly required. CAN COLLEGES AND UNIVERSITIES AFFORD TO IMPLEMENT HIGH TECH DISASTER PLANS? CAN THEY AFFORD NOT TO? Many higher education administrators are asking whether it is necessary for their institutions to embark on the same costly path of high technology contingency planning as the private sector. Are colleges and universities different from commercial organizations in their reliance on information technology? Although, in our estimation, educational institutions have more per capita computing power than most business corporations, we also believe that colleges and universities generally depend less on day-to-day (or hour-by-hour) computer operations. At a time of financial constraints and operational priorities forced by the threat of seceding user departments, it is no wonder that most administrators procrastinate when auditors or magazine articles urge contingency planning. An open-ended electronic mail survey of Canadian colleges and universities, conducted by McGill University[3] in 1989, concluded that most institutions were talking about contingency planning, that auditors were regularly mentioning the subject in their annual reports, and that directors of computing centers and services were having periodic pangs of conscience, but that few institutions had actually completed the task. In addition to the cost of high tech solutions, an often cited deterrent was the documentation that usually accompanies a disaster recovery plan. According to one university administrator, a professionally developed disaster recovery plan had weighed in with 500 pages, containing information of limited value. Why limited? Because a disaster recovery plan must be simple enough that it can be maintained and its procedures tested regularly, without a great expenditure of time and money. The challenge facing colleges and universities is how to protect information system resources without overstraining budgets and manpower. There is no point in buying such an expensive insurance policy that the premiums will kill you! MCGILL'S COMPROMISE APPROACH-- FOCUS ON MANAGEABLE THREATS This was the challenge facing McGill when, in the spring of 1991, the University embarked on a cautious path of analyzing the risks associated with computing and information technology. The project directives included the assessment of central and departmental, academic and administrative, local and remote computing, as well as of McGill's sprawling communication network. Finally, the task called for the development of a "metaplan," a term coined by the director of the computing center to define the requirement for a high-level master plan, i.e., a plan for the process of planning for disaster recovery/contingencies. The approach taken by McGill placed the emphasis on the first stage of contingency planning and on risk assessment, focusing on preventive measures. The ensuing study recognized the changing face of computing in a typical campus environment. Planning for disaster in isolation of user departments and without full recognition of the increasing significance of departmental computing would have rendered the exercise practically useless. McGill engaged the services of a consultant organization, well versed in the computing environment of educational institutions. Due to financial constraints, it was vital to structure the consultant's assignment for maximum effectiveness at a minimum cost. The study had to reflect McGill's environment accurately, while keeping overall costs under $15,000. The departments and individuals interviewed were carefully chosen with the aim of selecting a representative, but lean, cross-section of the University. It was felt that the use of a consultant signaled management commitment to the user community. The external consultant provided an independent and unbiased view and ensured that the risk assessment task was well structured and completed on target. McGill staff provided clerical assistance and supported the analysis process. The outside consultant was able to maintain an objective view of the diverse departmental computing scene, while working closely with the MIS department and the computing center. The effort was well received by the users and departmental managers. In fact, the challenge was to include everyone on the limited sample list who wished to be interviewed. McGill found that its vulnerability extended far beyond major disasters threatening the central computer operations. While the possibility of major disasters is always present, disasters of a lesser impact are more likely to occur. Often, these threats are easier to manage in terms of both contingency planning and preventive measures. ACADEMIC ENVIRONMENTS ARE DIFFERENT Some are proud of the difference, others are apologetic, but all agree that academic environments are different from commercial, industrial, and even government organizations. Recognizing academic idiosyncrasies, our analysis took an unorthodox approach to risk assessment, based on two fundamental factors: combining formal methodology with common sense, and recognizing that computing is where the user is. While the use of common sense was emphasized, formal methodology provided a structure to the investigation. The second key consideration reflected today's reality: as computing has become ubiquitous on campuses, physical and intellectual assets have become widely distributed, and their vulnerability is greater than that of the generally well guarded central computing facilities. Today, in most institutions, the term "distributed computing" applies to administrative as well as research and teaching systems. For example, at McGill, the accounting department manages its own network for optical document storage, the bookstore runs its own stand-alone system, the payroll department has developed an online front end to the University pay system, and the list goes on. Even though decentralized computing is an accepted norm in most institutions, auditors are still focusing on the large and easily visible facilities, as if these were the only areas at risk. Departmental computing facilities, local area networks, remote and distributed databases are rarely scrutinized by the auditors. Often, these facilities are managed by part-time managers, faculty members, or office staff with other key responsibilities. Many academic computing facilities are acquired through research grants which only cover the purchase price. As a result, departmental computing operations frequently lack the necessary funding for proper security and contingency measures. RISK ASSESSMENT The process of risk assessment is a usual precursor to contingency and disaster recovery planning. The logic flows like this: * "What are the assets of the organization?" * "Are these assets vulnerable? Are they exposed to any threats?" * "What is the likelihood that the assets will be destroyed?" * "Assets with the largest value and the greatest exposure will have the highest risk value, and these will have to be protected first," or, to use the formal definition: Risk Factor = Asset Value x Likelihood of Threat The McGill risk analysis process considered the University's unique character, the combination of central and distributed computing, and its decentralized management structure. The risks related to information systems were grouped into eighteen categories, common to several departments. In addition, several department-specific risk areas were identified. At the outset of the investigation, the following broad categories of assets were established, each broken down into subcategories: * Physical Assets/Environment * Computer and Communication Hardware * Software * Data * People and Services The assessment was based on interviews and documentation analysis of a broad cross-section of the University. Interviews were conducted with over forty people, representing twenty-six departments, laboratories, or faculties. In addition to the MIS department and the computing center, the analysis included medical and engineering research laboratories and teaching staff, as well as key administrative and many ancillary departments. SOME OF THE FINDINGS The principal findings of the study related to the physical exposure of assets (theft, equipment failure, fire, flood, power failure), the safeguard of information (equipment failure, security breaches, viruses), vulnerability through communication networks, and dependence on key individuals. The following is representative of the risk assessment findings. These, along with the risk exposures, were divided into two major categories: (1) those common to several departments of the University, and (2) those specific to one department or functional area. Common risks Among the common risks were destruction by fire, vandalism/loss through break-ins, poor location/environmental problems, inadequate backup procedures, and critical dependence on in-house applications. Destruction by fire Damage or destruction by fire is perhaps the most obvious risk affecting any organization. The increasing decentralization of information handling and information storage drastically increases the likelihood that some data, stored in electronic or hard copy format, will be destroyed by fire. The McGill campus enjoys many Ivy League characteristics, including older, stately buildings, which are an important part of its heritage. Only the most modern buildings are equipped with smoke or heat sensors, or sprinkler systems, and much of the campus has nothing. The University Administration Building houses most of the administrative departments, including accounting, admissions, human resources, management systems, payroll, the principal's office, purchasing, the registrar's office, and the secretariat. These offices handle and store vital information regarding the institution's day-to- day, as well as its long-term, operation. The study found that a considerable amount of the information was not available in electronic format and was not easily reproducible. The building was not equipped with a sprinkler system. Vandalism/loss through break-ins Even though McGill has not suffered losses through theft and vandalism to the same extent as many other institutions, the incidence is rising rapidly. At least some of the loss can be attributed to the false sense of security given by the University's earlier remarkably small losses. Most seriously affected have been laboratories on the ground floor of buildings, with entrances out of sight of security guards. The study found most computer laboratory equipment to be unsecured and often protected only by flimsy door locks. Push-button mechanical door locks had been installed which offered little deterrent even to the amateur thief; most of these locks could be opened in a matter of minutes by systematically trying different combinations. Poor location/environmental problems Poor location of computer equipment is often a risk factor. The inappropriate location of a computer laboratory exposes it to theft or vandalism; the location can also expose it to water or fire damage. The study found a high risk of loss or damage due to the location of several departmental computer installations. Inadequate backup procedures Adequate backup procedures are one of the most important safeguards against disaster. While backup procedures differ according to the importance of data, its frequency of change, volume, and many other factors, the need for protecting information through backup applies to all sizes and types of computers. The McGill study encountered a number of shortcomings in the backup procedures of departmental systems, e.g.: * backup tapes kept adjacent to the equipment without any off-site backup; * the last off-site backup tape six months old; * inadequate backup tape cycling; * full backups taken a month apart, and only the penultimate tape kept off-site (potential exposure: up to two months' work); * incremental backup accumulated on the same tape for a month, or longer (in this case the department risked losing up to a month's additions and modifications, not only in the event of a natural disaster, but even through a tape malfunction); * blind faith in the central computing facility (even though the central MIS department was carrying out rigorous procedures to ensure the protection of its users' information, users needed to be aware of these procedures and the extent of protection provided by the backups); * backup schedules arranged by department system managers to suit forgetful users who wanted to reload erased files (they relegated safety issues to second place); and * frequent failure to back up personal systems until some disaster struck. One of the significant benefits of the risk analysis study was an increase in user awareness and a recognition of the need for formal backup procedures. Critical dependence on in-house applications When we mentioned risk assessment to McGill personnel, the typical response was, "Oh, you mean fire, earthquakes, natural disasters?" The study found, however, that one of the University's greatest hidden risks was a dependence on individuals in critical positions who had either developed in-house applications or who were the sole custodian of vital information. While the adage "nobody is indispensable" is not disputed, the potential disruption caused by the sudden absence of a sole custodian of vital information could be catastrophic to the operation. The study found several applications which had "slipped" into critical roles in the institution's information flow, without formal documentation or the training of a backup person. Additional findings The final report, consisting of over seventy pages, contained several additional categories of risks such as illegal computer access, network security standards, exposure to fraud and computer viruses, the use of pirated software, and the need for informed software and hardware acquisition practices. Risks specific to individual departments Ask any network manager or departmental computer guru worth his or her salt, and you will hear some of the risks associated with their operation. The deputy manager or top users will add to the list, but they may also miss some of the obvious threats because they are too close to the operation. An assessment by a person detached from the everyday bustle will be able to quickly identify easy-to-correct situations. This was the case in the McGill study which, not surprisingly, found some loose ends in the computer center operation as well as in departmental systems. THE BOTTOM LINE--NEITHER HOT NOR COLD The study affirmed that a contingency plan for McGill must take into account the University's highly networked and distributed computing environment. A traditional cold site[4] backup plan would be costly (estimated at $50,000 to $100,000 initially and a similar amount annually), would be difficult to maintain, and would provide only partial backup for the complex operating environment. There was no evidence that McGill's needs warranted a hot site[5] backup (as would, for example, a financial institution with minute-by-minute reliance on its mainframe computers). It was clear, however, that good documentation must exist to allow McGill to recreate its central facilities, should a major disaster ever strike. The study also confirmed that McGill relies on its campus network as much as on its traditional mainframes, and that the network must be adequately protected within reasonable costs. Complete duplication of the fiber-based network was, however, judged to be impractical. Lastly, since much of McGill's exposures was found to be at a departmental level, it was concluded that the contingency planning process should also address decentralized applications and departmental networks. INITIAL IMPLEMENTATION AND THE WAY AHEAD As mentioned above, one of the purposes of the study was to develop a contingency "metaplan" for the University, to describe in detail the steps required should McGill decide to proceed with a full-blown contingency plan. The metaplan resulting from the study provided the option to develop a disaster recovery plan for the central computing facilities only, or to extend it for the entire institution. Six stages of the planning process were identified, together with an estimate of personnel days required for the implementation. The metaplan identified those tasks which explicitly require University personnel and those which may be carried out by either a consultant or McGill staff. The metaplan is flexible and places a considerable onus on user departments and groups to develop appropriate contingency measures under the guidance of the central organizations. As a first phase, the "easy to fix" items were identified and the responsible departments charged with implementing the solutions. In the second phase, the University will review the benefits of the costlier safeguards, such as sprinkler installation throughout the Administration Building. Finally, the implementation of the proposed contingency plan will be considered. Since the largest exposures were found at the departmental level, an important course of action was to increase local managements' awareness of risks. This involved articles in University publications, raising the issues and concerns with McGill management groups, and general consciousness raising. In the months since the report was presented, many departments have improved internal procedures and an increasing number of local systems managers are being hired to oversee local computing efforts. The dependence of the campus network and its vulnerability remain major issues. As indicated by the study, full protection against a major disaster is not possible; however, the concept of "selective parallelism" is being used to identify key areas where a modest number of redundant links and nodes can significantly increase the network's reliability and resiliency. CONCLUSIONS Most colleges and universities do not have the necessary resources to carry out and maintain an elaborate contingency plan for their information systems. Furthermore, even if a full-scale disaster recovery plan and hot site were to be established for the central computing facility, the burgeoning departmental systems would remain unprotected. In its recent study, McGill University found that exposure through inadequate backup of departmental systems and dependence on key players was far greater than from the traditional risk of a major disaster. McGill therefore concluded that it could live without a hot or a cold site, and instead decided to focus initially on those things which could be improved at little cost through management attention, e.g., better backups, less dependence on key people, and an awareness campaign across campus. We strongly recommend that if you have been thinking of developing a contingency and disaster recovery plan, but, like most, have been putting it off until funding improves or current crises subside, consider conducting a low-cost risk assessment that samples the whole campus. Concentrate on risk avoidance, and establish a simple contingency plan to protect against those risks which can be minimized through careful planning and user cooperation. ======================================================================== Footnotes: 1 Clifford Stoll, The Cuckoo's Egg (New York: Bantam Doubleday, 1989). 2 M. A. Snow, "Computers in Banking," Dealers' Digest Inc., April 1990, p. 22. In July 1989, the U. S. Federal Financial Institutions Examination Council (FFIEC) issued a statement regarding the need for a contingency and disaster recovery plan. Following this statement, the Federal Deposit Insurance Corporation (FDIC) issued a letter indicating that the implementation of a comprehensive plan is the responsibility of the banks' Board of Directors. 3 McGill University, located in Montreal, Quebec, is one of the oldest universities in Canada. A heavily research oriented university, McGill has a student enrollment of 19,545 full-time and 10,400 part-time, and 2,500 faculty. The University consists of a main campus downtown with some seventy buildings on eighty acres, as well as a 1,600-acre agricultural campus north of the city. 4 A cold site provides a computer room environment with suitable air conditioning, power, and telecommunications, but no actual computing equipment. 5 A hot site provides a fully configured computing system which can be turned on within a few hours of a disaster. ======================================================================== An Alternative to Costly High Tech Contingency Plans