Integrating Computing Resources: A Shared Distributed Architecture for Academics and Administrators Copyright 1994 CAUSE. From _CAUSE/EFFECT_ Volume 17, Number 2, Summer 1994. Permission to copy or disseminate all or part of this material is granted provided that the copies are not made or distributed for commercial advantage, the CAUSE copyright and its date appear, and notice is given that copying is by permission of CAUSE, the association for managing and using information resources in higher education. To disseminate otherwise, or to republish, requires written permission. For further information, contact Julia Rudy at CAUSE, 4840 Pearl East Circle, Suite 302E, Boulder, CO 80301 USA; 303-939-0308; e-mail: jrudy@CAUSE.colorado.edu INTEGRATING COMPUTING RESOURCES: A SHARED DISTRIBUTED ARCHITECTURE FOR ACADEMICS AND ADMINISTRATORS by Monica Beltrametti and Will English ABSTRACT: While capital funds for central computing at the University of Alberta were shrinking, campus departments were buying desktop computers and powerful workstations by the thousands. An institution-wide strategy was needed to leverage this rich distributed computing environment. Developing a shared distributed architecture has proved to be an effective response; managing the extended enterprise continues to be a challenge. At the University of Alberta we have defined a computing architecture that enables the University community to share computing resources distributed across campus. This architecture provides users with an integrated electronic environment, with easy and transparent access to academic and institutional data, high performance computing, printers, and other peripherals. The sharing stretches beyond departmental boundaries. On demand, researchers can, for example, share each others' workstations, and take advantage of idle cycles on servers that are otherwise managing institutional data. In this article, we explain how the vision of shared distributed computing was developed and how it is being implemented; provide a high- level description of the architecture and the users' view of the electronic environment; explain the technical, and especially the managerial, challenges behind and ahead of us; and list the campus-wide human infrastructures necessary to manage such an integrated electronic environment. Some historical background During the mid-80s, computing at the University of Alberta underwent some dramatic changes. As the economic boom of the oil industry declined, funds to the University shrank quickly. Consequently, capital funds for computing were reduced from $8 million to $1 million a year. This prevented the University's central computing organization from keeping up with technology. The reduction of funds happened at the same time that workstations became affordable. Researchers and University staff did not hesitate to take advantage of this new commodity and started realizing their own computing solutions using their own funds, such as research grants, or redirecting other sources of money. This led to a peculiar environment: at the end of the 80s the University computing organization was managing second or third-hand computing equipment, while the campus (with 7,000 staff members and 35,000 students) had bought state-of-the-art desktop computers by the thousands. In the early 90s, users started to perceive the burden of managing their own computing equipment and actively demanded that the central computing organization change in nature to provide support for distributed computing. The users also demanded that computing at the University be raised in profile to match the standards of computing that the University had enjoyed during the seventies, when funds were generous and computing at the University could compete with the top universities in North America. This is why in 1992 Computing and Network Services (CNS), the computing department at the University, formulated a strategic plan in collaboration with senior University administration and the user community in an attempt to define a road map that would restore the state-of-the-art computing and support that the University had enjoyed earlier. However, while we formulated the strategic plan, we knew that one important parameter had changed: the level of funding that had been enjoyed in the past would not be available again. The University of Alberta president recommended that the computing organization transfer salary funds to capital funds at a rate of $200,000 a year for five years, which would result in the department having an additional $1 million a year for capital allocation. Although painful to staff, this recommendation helped improve our financial position[1] The capital funds we could count on were still not large, and it was clear that raising the profile of computing despite inadequate funding meant that we had to develop a new way of thinking. Besides financial considerations, several other factors contributed to conceiving a new computing environment for the campus. Rationale for defining a shared distributed computing architecture The first shift in our thinking occurred when we started implementing the strategic plan. The plan contains fifteen strategies, which call for: (1) various aspects of distributed computing support, such as installation of a high-bandwidth campus network, services to help clients with the management of their own computing equipment, and the establishment of guidelines for purchasing equipment to ensure adequate support and interoperability ; (2) increased computing capacity for both researchers and administrators; (3) integration and modernization of administrative applications; and (4) enhanced support for instructional computing (instructional labs and tools for computer-assisted instruction). Originally, we distributed the strategies among the CNS management team with the mandate to prepare operational and financial plans for each strategy and then provide the resources and project management for implementation. We soon found that this was not the optimal way of implementing the strategic plan. Several difficulties arose. * The full force of wisdom and advice had not been sought from the user community at large. * There was a lack of uniformity and common agreement on the appropriate technology to be brought to the directions. * Where a technology solution was identified, service providers in different groups were implementing solutions that presupposed different levels of functionality and service to the user. Need to restructure central computing support No strategy could be implemented within the jurisdiction of one single manager. For example, integrating and modernizing administrative applications using a new client/server architecture could not be implemented by the Information Systems group alone, which in the past had primarily dealt with mainframes. Expertise was needed from staff with networking and workstation experience. Similarly, implementing a distributed open system environment for research could not be implemented by staff with only workstation background; expertise in system software and networking was required. We were very fortunate to be able to address this problem easily, thanks to some organizational changes that had occurred previously. The University of Alberta used to have two computing departments: Computing Services and the Office of Administrative Systems. The former department reported to the academic vice president, while the latter reported to the vice president for finance and administration. Following suggestions put forward as early as 1985 by some visionary University staff, in the spring of 1988 the two departments were put under a single director reporting to the academic vice president. At that time the two computing rooms were amalgamated; the rest of the organization, however, continued to operate as two independent entities. Administrators and researchers used different equipment and sought support from different staff. This led to substantial duplication of services, and in 1992 we carried the organizational change a step forward by eliminating in the computing organization the distinction between academic and administrative computing support, and reorganizing along technology lines (see Figure 1). Figure 1: Retrospective of computing reporting structure [FIGURE 1 NOT AVAILABLE IN ASCII TEXT VERSION] Even though the new organization was efficient in delivering day-to-day support to clients, it was still too rigid to implement the strategic plan effectively. However, thanks to a new way of thinking that had developed over the years, it was easy to break departmental barriers even further. We established task forces with staff from various CNS groups and technical backgrounds. Each task force has a leader and has the mandate to implement one strategy in the plan. By mixing staff with different backgrounds, new ideas soon started to flourish. Need for architectural cohesiveness and integrity It was good to generate new ideas, but we were left with the problem of organizing and selecting them to ensure effective and compatible implementation of the strategies defined in the plan. Filtering ideas was difficult, since within each task force we had staff with different technology backgrounds and beliefs: some were wedded to mainframes, while others had embraced workstation technology to the extreme. Also, due to the intricacy of distributed computing, some strategies had overlapping requirements. Since each group was working without an overall technological guideline, groups tended to find different technology solutions for the same requirements. For example, implementation of electronic forms and general-purpose communication required e-mail solutions. These solutions were being looked at separately, without an attempt at coordination. We therefore felt the need to define clear architectural directions for computing and networking, which could guide CNS and the users in their decisions. A Manager of Technology Directions was appointed, with the mandate of defining a computing and network architecture for the campus, taking advantage of as much staff wisdom as possible. To cause minimal disruption to the work of the task forces, he decided to formulate and publish the architecture plan one chapter at a time, starting with the issues that most urgently needed coordination between the task forces. Need for guiding principles Some basic principles were established as an underpinning to defining the campus computing and network architecture: * Do not resist the distributed computing trend. Clients had been frustrated for many years by the obsolescence of the central computing power and by the backlog of work in administrative applications development. As a result, many clients had already found alternative solutions using workstations. It was clear that distributed computing was happening no matter what. New solutions had to acknowledge this and also had to accommodate the desire of clients to choose their own solutions while still benefitting from a central coordinating infrastructure. * Avoid paying mainframe bills. Our budget was too tight to be able to think about mainframe upgrades. We had to find a way to increase and modernize the computational power using cheaper technology. * Capitalize on existing University investments. Our clients had already invested heavily in new technology by purchasing workstations by the hundreds. We had to find a way to capitalize on this investment. * Create an integrated electronic environment. It was clear that we could add value to the investments that already existed on campus by tying together the distributed computing resources- -for example, by providing easy access to academic and institutional data and to special devices such as high-speed printers or high- performance computers. Need to demolish barriers between academic and administrative computing systems As the architecture was being developed, an idea emerged that we think is fundamental to changing the way computing resources are used on campus. For decades academics and administrators used different computer systems. With the way technology has evolved lately, this is no longer necessary, nor is it desirable. A decade ago, when many scientists started moving away from aging mainframes, seeking faster response times from workstations, the manipulation of large amounts of institutional data could still only be handled by mainframes. Today, the same workstations being used for scientific work can also handle large databases. If used wisely, they can entirely replace mainframes for that use. Furthermore, in the past, UNIX was primarily used for scientific computing. Today UNIX, if used in connection with other tools, can also handle the secure manipulation of data. It seems that today many ingredients are available to define a uniform architecture for both academic and administrative computing. By making full use of campus computing resources, and pushed by the desire to make use of economies of scale, the CNS staff defined an architecture that allows academics and administrators to share resources across campus. Even though the sharing comes with considerable management challenges, there is strong support on campus to make the shared environment work. Description of shared distributed architecture A solid network is fundamental to a distributed architecture. In establishing this foundation, special attention must be placed on the differing communication and service needs of the administrator, the student, and the researcher. The infrastructure must, however, focus on the principles of sharing across these communities, while also operating in an open systems environment. Network backbone Distributed resources on campus are harnessed via a 64-kilometer, 100- million bits-per-second fiber-optic FDDI backbone network, which is installed along the University tunnels. The network is being connected to concentrators at preferred ring node locations. The concentrators are used to extend the fiber from each ring node location to the machine room of adjacent buildings. The fiber optic cable has twenty-four strands. Twelve strands have been reserved to implement the distributed architecture described below. The rest of the strands have been reserved for future applications such as energy and utility management. The network is being installed in phases. The north route is functioning, and the rest of the ring will be ready by the summer of 1994. The ring is being financed by converting salary funds into capital. Campus ring Two strands of the fiber optic backbone are used to form a general purpose "campus ring." The campus ring will harness most computers on campus (see Figure 2). This is done by connecting the fiber in the machine room of each building to routers, which in turn connect to the departmental local area networks (LANs). Figure 2: Campus and research rings [FIGURE 2 NOT AVAILABLE IN ASCII TEXT VERSION] The campus ring is being used to provide general purpose services, such as e-mail, file transfers, distributed printing, and the manipulation of institutional data. Institutional data will reside on six enterprise servers placed at strategic locations on campus and connected to the FDDI backbone network. To provide maximum performance, each enterprise server will be located near the local area network of the administrative unit that the server primarily supports. Information and applications of interest to only one administrative department will be handled within the department's local area network. The servers will handle information and applications that are of interest for all campus users. The servers will support enterprise-wide applications such as: * enterprise data used by many applications across the campus; * enterprise file libraries housing the programs and files needed to support the users in accessing the enterprise data; and * enterprise services supporting cross-platform applications, such as electronic University form processing or enterprise document management. Research ring Two strands of the backbone will be used to network research workstations across the campus. For each workstation, four fiber strands will be extended from the nearest concentrator to the workstation. The goal is to better utilize the campus investment in workstations by allowing researchers to execute their jobs on idle workstations. The ring can contain workstations and larger systems from different vendors. Workstations connected to the ring can be on researchers' desks or in instructional labs. Jobs submitted to the ring can be sequential or parallel: * Sequential program execution: Researchers will be able to utilize workstation cycles on the ring to execute their sequential jobs. Each workstation on the ring will run an NQS-like software package (Network Queuing System). The software allows users to submit a job to a queue on a particular machine on the ring or to submit a job to a "ring queue," which will then select the best available machine on the ring. * Parallel program execution: Researchers will be able to utilize workstations on the ring for parallel processing. The workstations will run the Parallel Virtual Machine (PVM) software package from Oakridge National Laboratory, a tool that provides an environment for designing, testing, and debugging parallel applications. Both software packages mentioned above allow the research ring to be formed at particular times of day. Researchers wishing to use their workstation alone can disconnect from the ring. They can reconnect to the ring when they desire. The campus ring and the research ring will be connected to allow researchers to have full use of the campus ring. We also anticipate that the enterprise servers will not be used at full capacity at all times. Researchers will be able to use the enterprise servers, especially at night, for the execution of their programs. There are two primary reasons for connecting: (1) If researchers need more computational power than their personal workstations can handle, they can connect to the ring to ship some of their programs to idle workstations on campus. At night, for example, computers in instructional labs can thus be easily used for production runs. (2) If researchers anticipate that their personal workstations will not be fully utilized for some time, they can "donate" some of their cycles to colleagues. "Foreign" programs running on a workstation are "time- sliced," so that researchers always have first call on their workstations. Many other institutions have already networked their workstations in this manner. The novelty here is that the workstations do not all reside in one machine room and are not centrally owned; they belong to researchers or instructional labs dispersed across campus. Obviously, there are various management and security issues that come along with the research ring. They are discussed below. CNS is also building a "farm" of workstations to be located in the CNS machine room. The farm will run the same tools as the research ring, and more. Researchers will be able to offload their programs to the farm as well, although here the intent is to provide the researcher a scale of machine that would not normally be at her or his desktop, or indeed to provide computing to researchers who do not have a machine at all. Open system environment The campus ring and the research ring will be managed under the auspices of the Open System Environment Project. This project is dealing with: * authentication services to ensure data security, such as Kerberos; * a user ID service that uniquely identifies campus users, such as UNIQUENAME; * a campus-wide network file system, such as AFS; * monitoring tools to measure computing resource usage in a distributed open system environment; and * system administration guidelines for managing distributed resources in a collaborative manner with faculties and departments. Some of above tools are already undergoing testing in CNS and will be expanded to the campus in the summer of 1994. This initiative has focused on pre-release beta versions of the Distributed Computing Environment (DCE) from IBM and its mutual operation with OSF-1 from Digital Equipment Corporation. It is through the services of DCE that data are being provided to the programs executing on the farm workstations, or across workstations connected to the campus ring and research ring networks. Users' view of an integrated electronic environment Many devices and utilities are already attached to the network backbone, and their number will grow over time. It is essential that we present to the user an easy interface to the network environment, hiding the complexity of the operation. The goal is to provide users (students, faculty, staff, and the community) access to an integrated electronic environment from any workstation on campus, be it located on the user's desk, in an instructional lab, or in the library. Access will be provided to academic data (library, special databases), to institutional data (student records, classes, forms, etc.), and to special devices (workstations, printers, publishing devices, high-performance computers). These resources might be situated on campus, directly attached to the backbone, or somewhere on the Internet. Migrating from mainframes The computing resources on the campus ring and the research ring will slowly replace our mainframes. We anticipate that by 1998 CNS will not host any mainframes. Replacing the central administrative applications with client/server applications is the greatest challenge to achieving this goal. We have already made good progress by establishing the preferred relational database management systems (RDBMS) foundation for the campus. Oracle was selected, subsequent to a detailed request for proposals. The new Oracle toolset was purchased in April 1992 and is currently installed in twenty-six distributed sites offering the capabilities of the RDBMS, CASE, and Oracle toolset and communication capabilities to the enterprise servers, as well as the RDBMS and communication capabilities on departmental servers. It was initially hoped that the current legacy databases could be accessed via the Oracle tools to provide the campus with easy accessibility to current data, as opposed to the archaic and cumbersome GIS reporting tool under IMS. This trial proved unsuccessful, and it was necessary to devise an alternative method. Via a selective and periodic extract, current legacy data are downloaded to the Oracle tables on some departmental servers, thus enabling campus users to access legacy data for query and reporting needs via Oracle tools. A project requiring longer development involves the transfer / redevelopment of current applications, and creation of new applications in the new distributed environment, with a goal of having the current mainframe and IMS facilities dispensed with before the end of the decade. During this phase, certain mainframe facilities such as document preparation, interactive training, librarian, and performance monitoring will be removed to help provide funding for the new client/server environment. A task force, led by the office of the vice president for finance and administration, is investigating whether the University can purchase industrial application packages running on Oracle, or whether we have to redevelop the applications ourselves. The task force has members from the major administrative departments, from faculty, and from CNS. A request for proposals was issued in November 1993, and the tenders are now being evaluated. A recommendation to the University will be forthcoming by the autumn of 1994. While reviewing industrial packages, the task force has looked at the existing University administrative processes. This has to be done using an iterative approach, as industrial packages have administrative processes already embedded in them. It is important that compromises be made wherever possible, so that we give the industrial packages a good chance of acceptance. We all agree that the University cannot afford to redevelop all applications. Supporting human infrastructures-- managing an extended enterprise Implementing and supporting a distributed computing environment comes with various managerial challenges. We believe these challenges far outweigh the technical challenges. In the past, central computing organizations dictated computing solutions to the campus. Today, both computing resources and computing expertise have spread on campus and management of a computing infrastructure has become an endeavor to be shared with the users. The University has addressed these management challenges by creating additional advisory committees as well as a CNS outreach program (see Figure 3). Figure 3: Managing an extended enterprise [FIGURE 3 NOT AVAILABL EIN ASCII TEXT VERSION] Advisory committees There are various tasks that we are trying to share and coordinate with user input and participation. Besides the University Computing Advisory Group (UCAG), which reports to the vice president for student and academic services, and which has existed for many years, we have established other advisory committees. Officially these committees report to the same vice president, but it is CNS that is managing them. All these committees have representatives from faculty and administration. Some of the committees also have student representatives. The committees help gather user requests, provide a sounding board for CNS plans, help establish policies and procedures, and help us coordinate campus activities related to information technology. * Network Advisory Committee The committee was used as a sounding board for the FDDI network backbone architecture. Current challenges comprise educating the campus on how to plan for and connect departmental LANs to the network, managing a presidential fund to cover the cost of connecting departmental LANs to the network, and teaching users about Internet access and potentials. * Numerically Intensive Computing Committee The committee advises on the directions for high-performance research computing on and off campus. The challenge ahead is management of the research ring. A special user committee will be set up to establish policies and procedures and to coordinate collaboration. Other challenges relate to connecting the research ring to high-performance computers in Alberta and elsewhere in North America. * Technology and Standards Advisory Committee The committee is used as a sounding board for the information technology directions formulated by CNS. It is also being used to help determine which information technology tools are recommended to the campus (and thus supported by CNS). * Information Systems Advisory Committee The committee advises on information systems directions and plans. It will be pivotal in helping us to migrate institutional data from the MVS mainframe to Oracle running on the enterprise servers and client workstations distributed across campus. Tasks ahead of us, for which we will seek advice from the committee are: --establishment of a training and support program for customers to learn about the new Oracle tools; --establishment of criteria to select some pilot projects to test the new Oracle tools; --establishment of University-wide criteria for selection of industry applications to run with Oracle, such as financial or student record systems; and --definition of a data dictionary and the data management infrastructure, deciding, for example, which data are to reside on the enterprise servers and which data are to be stored within departmental LANs. CNS outreach program Besides advisory committees, CNS has established additional outreach programs. * Campus user support group Departments with large computer needs have appointed local computing support staff. The primary task of these staff is to manage the departmental LANs. In the past, these departmental computing support staff have worked in isolation. CNS is now starting a program to establish an extended user support enterprise on campus. By exchanging information and coordinating training and user documentation, CNS will collaborate with the departmental support staff to provide computing support on campus. * Service representatives CNS has also appointed staff to help faculties and departments plan for computing and networking. The task of a service rep, a CNS staff member assigned to a faculty unit, is to understand the way computers are used in that unit, guide users in high-level plans, disseminate information on CNS directions, and keep CNS appraised of issues and user wishes. * Product specialists While service representatives are assigned to faculties, product specialists help users with the use of specific products. Since different faculties may in many cases use the same products, product specialists transcend faculty and departmental boundaries. Product specialists keep track of departmental expertise and experience and establish a network among users with similar technology interests. * Help desk The help desk collaborates closely with the campus computing support group, the service representatives, and the product specialists to establish and maintain a network of users with various expertise and responsibilities in information technology. Computing and Network Services is still adapting our planning and management style to our new shared distributed architecture. Initial feedback from the user community is good, but we feel there is still room for improving how we capitalize on that feedback and how we use it to constantly improve our services. ======================================================================== Footnote: 1 See "Computing Services Planning, Downsizing, and Re-organization at the University of Alberta," CAUSE/EFFECT, Fall 1993, for a full discussion of the University of Alberta's strategic planning and organizational downsizing experience. ======================================================================== Monica Beltrametti is curently Director of Business Relations with the Rank Xerox Research Centre in Meylan, France. Until the summer of 1993, she was Director of Computing and Network Services at the University of Alberta, where she led the planning, downsizing, and reorganization process discussed in this article. Before joining the U of A, Dr. Beltrametti was Director of Software Development with Myrias Research Corporation, developer of parallel computers. She earned her doctorate in astrophysics from Ludwig Maximilian University in Munich, Germany. Will English assumed the role of Director of Computing and Network Services at the University of Alberta in July 1993. His career with the University has spanned two decades and included such computing roles as Manager of Production Support, Manager of Data Communications and Networks, and Manager of Technology Directions. Prior to his work with higher education from 1964 onwards, he was with Sperry Univac in major systems technical and marketing support to government, education, and industry. ************************************************************************ Integrating Computing Resources: A Shared Distributed Architecture for Academics and Administrators