Integrating Computing Resources: A Shared Distributed Architecture for Academics and Administrators Copyright CAUSE 1994. This paper was presented at the 1993 CAUSE Annual Conference held in San Diego, California, December 7-10, and is part of the conference proceedings published by CAUSE. Permission to copy or disseminate all or part of this material is granted provided that the copies are not made or distributed for commercial advantage, that the CAUSE copyright notice and the title and authors of the publication and its date appear, and that notice is given that copying is by permission of CAUSE, the association for managing and using information technology in higher education. To copy or disseminate otherwise, or to republish in any form, requires written permission from CAUSE. For further information: CAUSE, 4840 Pearl East Circle, Suite 302E, Boulder, CO 80301; 303449-4430; e-mail info@cause.colorado.edu Integrating Computing Resources: A Shared Distributed Architecture for Academics and Administrators Dr. Monica Beltrametti, Director Computing and Network Services University of Alberta Edmonton Alberta Abstract At the University of Alberta we have defined a computing architecture that enables the University community to share computing resources distributed across campus. This architecture provides users with an integrated electronic environment, with easy and transparent access to academic and institutional data, high performance computing, printers, and other peripherals. The sharing stretches beyond departmental boundaries. On demand, researchers can, for example, share each others' workstations, and take advantage of idle cycles on servers that are otherwise managing institutional data. We explain how the vision of shared distributed computing was developed and how it is being implemented; we give a high-level description of the architecture and the users' view of the electronic environment; we show the budgetary advantages of the project; we explain the technical, and especially the managerial challenges behind and ahead of us; and we list the campus-wide human infrastructures necessary to manage such an integrated electronic environment. Some Historical Background During the mid-eighties computing at the University of Alberta underwent some dramatic changes. As the economic boom of the oil industry declined, funds to the University shrank quickly. Consequently capital funds for computing were reduced from $8 to $1 million a year. This prevented the University's central computing organization from keeping up with technology. The reduction of funds happened at the same time as workstations became affordable. Researchers and University staff did not hesitate to take advantage of this new commodity and started realizing their own computing solutions using their own funds,such as research grants or re-directing other sources of money. This led to a peculiar environment: at the end of the eighties the University computing organization was managing second or third- hand computing equipment, while the campus (having 7,000 staff members and 35,000 students) had bought state-of-the-art desktop computers by the thousands (see Table 1). In the early nineties, users started to perceive the burden of managing their own computing equipment and actively demanded that the central computing organization change in nature to provide support for distributed computing. The users also demanded that computing at the University beraised in profile to match the standards of computing that the University had enjoyed during the seventies when funds were generous and computing at the University could compete with the top universities in North America. This is why in 1991 Computing and Network Services (CNS), the computing department at the University of Alberta, formulated, in collaboration with senior University administration and the user community, a strategic plan in an attempt to define a road map that would restore the state-of-the-art computing and support that the University had enjoyed earlier (see "Computing Services Planning, Downsizing, and Re-organization at the University of Alberta" proceedings of the CAUSE conference,Dallas, December 1992; and "Networking Computers and People on Campus and Beyond" Computing and Network Services strategic plan, University of Alberta, February 1992,). While we formulated the strategic plan we knew, however, that one important parameter had changed: the level of funding that had been enjoyed in the past would not be available again. The University President, Dr. Paul Davenport, recommended that the computing organization transfer salary funds to capital funds at a rate of $200,000 a year for five years, which would result in the department having an additional $1 million a year for capital allocation. Although painful to staff, this recommendation helped improve our financial position. The capital funds we could count on were, however, still not large and it was clear that raising the profile of computing despite this meant that we had to develop a new way of thinking. Besides financial considerations, several other factors contributed to conceiving a new computing environment for the campus. Figure 1:(FIGURE NOT AVAILABLE) Rationale for Defining a Shared Distributed Computing Architecture Implementing the Strategic Plan The first shift in our thinking occurred when we started implementing the strategic plan. The plan contains fifteen strategies. They call for: *various aspects of distributed computing support, such as installation of a high-bandwidth campus network, services to help clients with the management of their own computing equipment, and the establishment of guidelines for purchasing equipment to ensure adequate support and interoperability * increased computing capacity for both researchers and administrators *integration and modernization of administrative applications *enhanced support for instructional computing (instructional labs and computer-assisted instruction tools). Originally, we distributed the strategies among the CNS management team with the mandate to prepare operational and financial plans for each strategy and then provide the resources and project management for implementation. We soon found out that this was not the optimal way of implementing the strategic plan. Several difficulties arose. Need to Abolish the Distinction between Administrative and Academic Computing Support Staff No strategy could be implemented within the jurisdiction of one single manager. For example, integrating and modernizing administrative applications using a new client/server architecture could not be implemented by the Information Systems group alone, which in the past had primarily dealt with mainframes. Expertise was needed from staff with networking and workstation experience. Similarly, implementing a distributed open system environment for research could not be implemented by staff with only workstation background; expertise in system software and networking was required. We were very fortunate to be able to address this problem easily, thanks to some organizational changes that had occurred previously. Figure 2: (FIGURE NOT AVAILABLE) The University of Alberta used to have two computing departments: the Office of Administrative Systems and Computing Services. The former department reported to the V.P. Finance and Administration and the latter to the V.P. Academic. Following suggestions put forward as early as 1985 by some visionary University staff, in the spring 1988 the two departments were put under a single director reporting to the V.P. Academic. At that time the two computing rooms were amalgamated; the rest of the organization, however, continued to operate as two independent entities. Administrators and researchers used different equipment and sought support from different staff. This led to substantial duplication of services and in 1992 we carried the organizational change a step forward by eliminating in the computing organization the distinction between academic and administrative computing, reorganizing along technology lines. Even though the new organization was efficient to give day-to-day support to clients, it was still too rigid to implement the strategic plan effectively. Thanks, however, to a new way of thinking that had developed over the years it was easy to break departmental barriers even further. We established task forces with staff from various CNS groups and technical backgrounds. Each task force has a leader and has the mandate to implement one strategy in the plan. By mixing staff with different backgrounds new ideas soon started to flourish. Need for Architectural Cohesiveness and Integrity It was good to generate new ideas, but we were left with the problem of organizing and selecting them to ensure effective and compatible implementation of the strategies defined in the plan. Filtering ideas was difficult since within each task force we had staff with different technology backgrounds and beliefs: some were wedded to mainframes, while others had embraced workstation technology to the extreme. Also, due to the intricacy of distributed computing some strategies had overlapping requirements. Since each group was working without an overall technological guideline, groups tended to find different technology solutions for the same requirements. For example, implementation of electronic forms and general purpose communication required e-mail solutions. These solutions were being looked at separately without an attempt to coordinate. We therefore felt the need to define clear architectural directions for computing and networking, which could guide CNS and the users in their decisions. Will English, who had helped pinpoint these problems, was appointed manager, Technology Directions, with the mandate to define a computing and network architecture for the campus, taking advantage of as much staff wisdom as possible. To cause minimal disruption to the work of the task forces, he decided to formulate and publish the architecture plan a chapter at a time, starting with the most urgent issues that needed coordination between the task forces. Guiding Principles The campus computing and network architecture is being defined with some basic principles in mind: * Do Not Resist the Distributed Computing Trend: Clients had been frustrated for many years by the obsolescence of the central computing power and by the backlog of work in administrative applications development. As a result many clients had already found alternative solutions using workstations. It was clear that distributed computing was happening no matter what. New solutions had to acknowledge this and had to accommodate the desire of clients to choose their own solutions while still benefiting from a central coordinating infrastructure. * Avoid Paying Mainframe Bills: Our budget was too tight to be able to think about mainframe upgrades. We had to find a way to increase and modernize the computational power using cheaper technology. *Capitalize on Existing University Investments: Our clients had already invested heavily in new technology by purchasing workstations by the hundreds. We had to find a way to capitalize on this investment. *Create an Integrated Electronic Environment: It was clear that we could add value to the investments that already existed on campus by tying together the distributed computing resources; for example, by providing easy access to academic and institutional data and to special devices such as high speed printers or high performance computers. Breaking the Barriers between Academic and Administrative ComputingAs the architecture was being developed an idea emerged which we think is fundamental to changing the way computing resources are used on campus. For decades academics and administrators used different computer systems. With the way technology has evolved lately, this is no longer necessary, nor desirable. A decade ago when many scientists started moving away from aging mainframes, seeking faster respond times from workstations, manipulation of large amounts of institutional data could still only be handled by mainframes. Today, the same workstations being used for scientific work can also handle large data bases. If used wisely, they can entirely substitute mainframes for that use. Furthermore, in the past, UNIX was primarily used for scientific computing. Today UNIX, if used in connection with other tools, can also handle secure manipulation of data. It seems that today many ingredients are available to define a uniform architecture for both academic and administrative computing. Pushed by the desire to make use of economies of scale and by making full use of campus computing resources, the CNS staff defined an architecture that allows academics and administrators to share resources across campus. Even though the sharing comes with considerable management challenges, there is strong support on campus to make the shared environment work. We describe next this shared distributed architecture. Description of the U of A Shared Distributed Architecture The Network Backbone Fundamental to a distributed architecture is a solid network. Distributed resources on campus are harnessed via a fast network backbone. The backbone being installed is a 100 million bits per second fibre optic FDDI network installed along the University tunnels in a figure of eight (see Figure 3). The network is being connected to concentrators at preferred ring node locations (denoted with letters in the figure). The concentrators are used to extend the fibre from each ring node location to the machine room of adjacent buildings. The fibre optic cable has 24 strands. Twelve strands have been reserved to implement the distributed architecture described below. The rest of the strands have been reserved for future applications, such as energy and utility management.The network is being installed in phases. The north route is functioning, the rest of the ring will be ready by 1994. The ring is being financed by converting salary funds into capital. Figure 3: (FIGURE NOT AVAILABLE) The Campus Ring Two strands of the fibre optic backbone are used to form a general purpose Campus Ring. The Campus Ring will harness most computers on campus. This is done by connecting the fibre in the machine room of each building to routers which, in turn, connect to the departmental local area networks (LANs) (see Figure 4). The Campus Ring is being used to provide general purpose services, such as e-mail, file transfers, distributed printing, and manipulation of institutional data. Institutional data will reside on six enterprise severs placed at strategic locations on campus and connected to the FDDI backbone network. To provide maximum performance, each Enterprise Server will be located near the local area network of the administrative unit that the server primarily supports. Information and applications of interest to only one administrative department will be handled within the department's local area network. The servers will handle information and applications that are of interest for all campus users. The Servers will support enterprise-wide applications, such as: * enterprise data used by many applications across the campus * enterprise file libraries housing the programs and files needed to support the users in accessing the enterprise data * enterprise services supporting cross-platform applications, such as electronic University form processing or enterprise document management. We are currently installing new information system tools which will eventually handle this client/server architecture. Some of the tools will have to evolve further before they will be able to handle our desired environment. Figure 4:(FIGURE NOT AVAILABLE) Figure 5:(FIGURE NOT AVAILABLE) Research Ring Two strands of the backbone will be used to network research workstations across campus. For each workstation, four fibre strands will be extended from the nearest concentrator to the workstation. The goal is to better utilize the campus investment in workstations by allowing researchers to executetheir jobs on idle workstations. The ring can contain workstations and larger systems from different vendors. The workstations connected to the ring can be on researchers' desks or in instructional labs. The jobs submitted to the ring can be sequential or parallel: * Sequential Program Execution: Researchers will be able to utilize workstation cycles on the ring to execute their sequential jobs. Each workstation on the ring will run an NQS-like software package (Network Queuing System). The software allows users to submit a job so a queue on a particular machine on the ring or to submit a job to a "ring queue" which will then select the best available machine on the ring. *Parallel Program Execution: Researchers will be able to utilize workstations on the ring for parallel processing. The workstations will run the Parallel Virtual Machine (PVM) software package from Oakridge National Laboratory, a tool which provides an environment for designing, testing, and debugging parallel applications. Both software packages mentioned above allow the research ring to be formed at particular times of day. Researchers wishing to use their workstation alone can disconnect from the ring. They can connect to the ring when they desire. There are two primary reasons for connecting: *if researchers need more computational power than their personal workstations can handle, they can connect to the ring to ship some of their programs to idle workstations on campus. At night, for example, computers in instructional labs can thus be easily used for production runs. *if researchers anticipate that their personal workstations will not be fully utilized for some time, they can donate some of their cycles to colleagues. "Foreign" programs run on a workstation are always "niced," so that researchers always have first call on their workstations. Many other institutions have already networked workstations in this manner. The novelty here, is that the workstations do not all reside in one machine room and are not centrally owned; they belong to researchers or instructional labs dispersed across campus. Obviously, there are various management and security issues that come along with the research ring. We describe them later in the paper. Sharing of Resources The Campus Ring and the Research Ring will be connected to allow researchers to have full usage of the Campus Ring. We also anticipate that the Enterprise Servers will not be used at full capacity all the time. Researchers will be able to use the Enterprise Servers, especially at night, for the execution of their programs. CNS Workstation Farm CNS is also building a "farm" of workstations to be located in CNS machine room. The farm will run the same tools as the Research Ring, and more. Researchers will be able to off-load their programs to the farm as well. Open System Environment The Campus Ring and the Research Ring will be managed under the auspices of the Open System Environment Project. This project is dealing with: * authentication services to ensure data security, such a Kerberos *a user ID service that uniquely identifies campus users, such as UNIQUENAME * a campus-wide network file system, such as AFS * monitoring tools to measure computing resource usage in a distributed open system environment *system administration guidelines for managing distributed resources in a collaborative manner with faculties and departments The Demise of Mainframes The computing resources on the Campus Ring and the Research Ring will slowly replace our mainframes. We anticipate that by 1998 CNS will not host any mainframes. The Users' View of an Integrated Electronic Environment Many devices and utilities are already attached to the network backbone and their number will grow over time. It is essential that we present to the user an easy interface to the network environment, hiding from the user the complexity of the operation. The goal is to provide to users (students, faculty, staff, and the community) access to an integrated electronic environment from any workstation on campus, be it located on the user's desk, in an instructional lab, or in the library. Access will be provided to academic data (library, special data bases), to institutional data (student records, classes, forms, etc.), and to special devices (workstations, printers, publishing devices, high performance computers). These resources might be situated on campus, directly attached to the backbone, or they might be somewhere on the Internet. Figure 6: (FIGURE NOT AVAILABLE) Budgetary Gymnastics to Move from Mainframes to a Distributed Architecture The shared distributed computing environment we want to establish will be much cheaper than the current mainframe environment (see first graph in Figure 7). However, while mainframe maintenance constitutes a large drain on our budget, mainframes are also responsible for very substantial revenues for CNS (see second graph in Figure 7). Figure 7: (FIGURE NOT AVAILABLE) As the mainframes disappear, the revenue will disappear with them. We have agreed with senior University administration that we would not seek other sources of revenue, once the mainframes are decommissioned. It was felt that internal University accounting is expensive, and that computing should eventually be regarded as a utility. We all believe this is a very desirable environment to establish for users, but it leaves us with an interesting budgetary challenge: *Users will move off mainframes faster than we can transfer mainframe applications (such as administrative applications) to a distributed client/server environment. Therefore, as users move off mainframes, the decline in revenues will be faster than the decline in mainframe maintenance costs. The result is a net loss in CNS revenue. * During the transition period from central to distributed computing, we will have to maintain both types of systems. Our expenses will therefore temporarily grow. The last graph in Figure 7 shows the overall projected operating costs of moving from mainframes to the shared distributed computing environment. We are currently devising with senior University administration various methods of how to bridge the transition period. We are extremely fortunate to have a senior University administration that understands the issues, and is supportive of initiatives (such as ours) aimed at improving University operations. Supporting Human Infrastructures - Managing an Extended Enterprise Implementing and supporting a distributed computing environment comes with various managerial challenges. We believe that these challenges far outweigh the technical challenges. In the past, central computing organizations dictated computing solutions to the campus. Today, both computing resources and computing expertise have spread on campus and management of a computing infrastructure has become an endeavor to be shared with the users. Advisory Committees There a various tasks which we are trying to share and coordinate with user input and participation. Besides the University Computing Advisory Group (UCAG) which reports to the V.P., Student and Academic Services, and which has existed for many years, we have established other advisory committees. Officially these committees report to the same V.P., but it is CNS that is managing them. All these committees have representatives from faculty and administration. Some of the committees also have student representatives. The committees help gather user requests, provide a sounding board for CNS plans, help establish policies and procedures, and help us coordinate campus activities related to information technology: *Network Advisory Committee (NAC): The committee was used as a sounding board for the FDDI network backbone architecture. Current challenges comprise educating the campus on how to plan for and connect departmental LANs to the network; managing a presidential fund to cover the cost of connecting departmental LANs to the network; and teaching users on Internet access and potentials. *Numerically Intensive Computing Committee (NICC): The committee advises on directions for high-performance research computing on and off campus. The challenge ahead is management of the Research Ring. A special user committee will be set up to establish policies and procedures and to coordinate collaboration. Other challenges relate to connecting the Research Ring to high- performance computers in Alberta and elsewhere in North America. * Technology and Standards Committee (TASC): The committee is used as a sounding board for the information technology directions formulated by CNS. It is also being used to help determine which information technology tools are recommended to the campus (and thus supported by CNS). *Information Systems Advisory Committee (ISAC): The committee advises on information systems directions and plans. The committee will be pivotal to help us migrate institutional data from the MVS mainframe to Oracle running on the enterprise servers and client workstations distributed across campus. Tasks ahead of us for which we will seek advice from the committee are: *establishment of a training and support program for customers to learn about the new Oracle tools * establishment of criteria to select some pilot projects to test the new Oracle tools * establishment of University-wide criteria for selection of industry applications to run with Oracle, such as financial or student record systems. *CNS will define a data dictionary and the data management infrastructure, deciding, for example, which data is to reside on the enterprise servers and which data is to be stored within departmental LANs. Figure 8: (FIGURE NOT AVAILABLE) CNS Outreach Program Besides advisory groups, CNS has established additional outreach programs: *Campus Computing Support Group: Departments with large computer needs have appointed local computing support staff. The primary task of these staff is to manage the departmental LANs. In the past, this departmental computing support staff has worked in isolation. CNS is now starting a program to establish an extended user support enterprise on campus. By exchanging information, coordinating training and user documentation, CNS will collaborate with the departmental support staff to provide computing support on campus. * Service Representatives: CNS has also appointed staff to help faculties and departments plan for computing and networking. The task of a Service Rep, a CNS staff member assigned to a faculty, is to understand the way computers are used in that unit, guide users in high-level plans, disseminate information on CNS directions, and keep CNS appraised of issues and user wishes. *Product Support Specialists: While Service Representatives are assigned to faculties, Product Specialists help users with the use of specific products. Since different faculties may in many cases use the same products, Product Specialists transcend faculty and departmental boundaries. Product Specialists keep track of departmental expertise and experience and establish a network among users with similar technology interests. *Help Desk: The Help Desk collaborates closely with the Campus Computing Support Group, the Service Representatives, and the Product Specialists to establish and maintain a network of users with various expertise and responsibilities in information technology. CNS is still learning how to manage these outreach programs. Initial feedback from the user community is good, but we feel that there is still room for improving how we capitalize on the information we gather from the user community and how we use it to constantly improve services.