A critical aspect of cloud-based operations that can receive less attention than it deserves in early planning is structuring how to recover normal operations during an outage or other disruption.
Disaster recovery has always been the red-headed stepchild of any moves to upgrade business computer systems. Particularly if an enterprise has never had a serious outage or incidence of downtime, it can be all too easy to relegate planning for such occurrences until "after we figure out the important stuff," like how basic organizational work gets done under a new system. When an enterprise plans to move to a cloud environment or modify an existing cloud structure, thinking about backup and recovery tasks can get mentally relegated to the "it's just a group of maintenance tasks that we already have a system for" bucket. It's a comfortable trap to fall into.
Hiring an outside cloud disaster recovery (CDR) expert is a logical first step to address backup and CDR concerns if the enterprise has the budget. However, this choice carries with it the risk of whatever expert is selected having an existing prejudice in key decisions such as the type of cloud infrastructure to use, a preferred list of CDO models or planning sequences, or a bias toward particular software products, security practices, and other solutions that may not be the best fit for a client's business type.
Another common mistake is to assume workflows in a cloud environment will be identical to those of a standalone environment. Even if the goal is to leave workflows unchanged for less disruption of employee duties in the changed environment, rethinking how work gets done is a worthwhile periodic exercise in any business setting. These concerns mean that, before picking up the phone to call an expert, there needs to be a plan for what happens before formulating an actual cloud disaster recovery (CDR) plan.
Preplanning for CDR
Even if the organization has undergone something similar, a full-blown analysis of how workflows will proceed under the new cloud structure is needed. This calls for listing all the tasks employees must follow to carry out each primary task of handling routine business activities. This process can be as complicated as redesigning the enterprise workload.
Workload redesign starts with defining existing user and system flows. If these processes are already documented, that's already a step forward. However, suppose that the documentation is not up to date. In that case, that deficit needs to be rectified to help redesign how work should flow under the new system, which presumably will incorporate changes. User flows outline how users move through applications and involve user interface design, user interactions with software and systems, decision points, and all steps required to accomplish all major tasks. This process maps a user-centric view of the overall system and provides a blueprint for how it should work in its next incarnation. System flows document workload actions such as input and output processing, how data moves through the system, and the interactions between workload elements, external APIs, and backend services. Accomplishing this cataloging of workflows requires the analyst to map out all workload processes and user interactions within every part of the intended system. It will throw into sharp relief the most crucial flows.
With that information in hand, the next step is designing how the workload should function under the new system. This process includes several essential steps. First, all critical stakeholders, including users, technical workers, and business analysts, will be interviewed to help outline user interactions and workload dependencies under the new system. Second, diagraming how well the new system meets the requirements of the superseded one. Third, observing and flowcharting the present workload, necessary employee interactions, and communications between parts. This includes reviewing systema and user-activity logs, performance metrics, and other environmental factors to inform the design of new workload procedures.
Next is listing all the workload flows as they should operate under the new system, categorizing them between user and system flows, defining the start and end points of each workflow, documenting all user interactions, and in the case of system flows, also identifying underlying catalysts and each flow's expected outcomes. Diagram each flow step by step, showing how actions and processes can be triggered at each point. This will help clarify how each flow affects user experiences and the overall workload, and provide flowcharts that make it less likely to overlook any critical step in any process. Understand also that this charting process will be a permanent work-in-progress for the new system, particularly if it was not under the old system. As workload flows are more closely defined or completely redesigned, stakeholders must update and review the diagrams and descriptions for completeness and accuracy.
Finally, document the business processes for each flow and map their intended connection to each process. This can explain the improvements the new system offers compared to its predecessor and highlight the relevance of each flow to overall enterprise goals. Be sure to map all the flows to each business process that each flow supports, because there will likely be some overlap. This plan will help simplify and guide the construction of the new system.
Choose the Right CDR Model
Another foundational decision about a new CDR system is selecting the CDR system model that best suits the enterprise's business model and finances. There are four basic models.
The simplest is the traditional model, usually called "backup and restore," which involves doing daily data backups. Traditionally this backup has been done to magnetic media that is physically stored offsite, but now it's most often done via the internet to a storage facility far enough away from an enterprise's place of business to avoid the possibility of a technical (e.g., power outage) or environmental disaster (e.g., fire, hurricane, earthquake) widespread enough to affect user locations and storage locations at the same time. The chief advantage of this method is that it's usually the lowest cost. The significant disadvantage is that restoring a data center from backup is time-consuming and can't guarantee continuous operations if that is important.
"Pilot Light," the next least expensive alternative, calls for a small-scale version of the operations environment to run in parallel at a separate, cloud-accessible location. In the event of need, this mini environment can be ramped up in a few hours to restore an alternative way to process data and financial transactions. Costs are moderate compared to the two most comprehensive CDR models below. However, the time delay is still a downside consideration unless the primary site goes down at a time when normal business operations are low anyway.
"Warm Standby" is the next model, under which a functional but not completely ramped-up version of the enterprise's processing environment is maintained at a remote but cloud-available location. Costs are higher than for the pilot light alternative, and the backup site may not handle enterprise activity at scale for some variable period. However, priority operations can be maintained with little or no downtime, depending on the structure of the potential failover process.
"Multi-Site" is the costliest but most complete alternative. Under this plan, the enterprise maintains one or more complete and parallel processing sites at dispersed geographic locations. Although it's the most expensive option, this alternative usually offers nearly instantaneous and transparent failover services. It's a must for enterprises that conduct business or internal processing on a 24-hour basis.
Regardless of the enterprise's CDR model, there should also be a backup plan for the plan itself. For example, the "backup and restore" option should be used as a secondary fallback. The purpose here is to provide some insurance against such problems as automated data replication unintentionally passing malware of any kind to a secondary processing site or "trusted" data archive, or user error causing deletion of data or transactions before potential replication takes place. Also, backup tapes in an off-site location rather than replicated disk drive images stored within the shop can provide the added protection of an air gap between versions of an operating environment. (For example, some ransomware incidents have included hacker-attempted deletion of backups stored elsewhere in the cloud.) Although having to fall back to offsite backup tapes will interfere with ideal recovery time objectives made possible by other CDR options, uncorrupted offsite backup tapes may be superior to having to track down and eliminate malware because there's no alternative.
Formulating the Basic Plan and Beyond
All these admittedly time-consuming steps are needed to make the best choices for a basic CDR strategy. With the information gained from the above steps, enterprises can feel ready to approach cloud service providers (CSP) and make intelligent decisions about which CSP's service offerings will best fit their environment. Many other choices must be made, particularly if an enterprise plans to set up its cloud environment. Still, many of them can be paused briefly while reviewing CSP offerings, because most CSPs have their preferred ways of handling those aspects.
For example, moving data across the cloud requires security measures. Data in motion and at rest, as well as in live environments, should all be encrypted. Encryption protocols should be updated periodically to counteract the constantly morphing threat of data theft or corruption. All procedures to achieve a failover in the event of disaster should be automated, documented on paper or another medium less likely to be affected by environmental problems, and practiced regularly. Both functions require planning before a CDP plan can safely be adopted.
The documentation must be updated whenever there's a change to procedures, and that updating must be someone's specific responsibility. The documentation must include details of roles and responsibilities and contact information for personnel vital to managing a failover. Employees must be trained in how to access this documentation, what to do in what order to react to a disaster, and how to monitor the environment, preferably with a set of cloud-environment monitoring tools. Any changes in responsible personnel should require a retesting of CDR procedures whenever that occurs.
Don't forget to check how much storage capacity will be needed for the new system. How much capacity is in use for the old system won't be sufficient. As time passes, storage needs gradually increase. Determining how much storage has increased for the past year will provide some guidance, but budgetary considerations will also play a role.
A DR Plan for Certain Cloud Services
DR for third-party cloud services practically requires its business plans. Unfortunately, some tools offered for cloud management — such as Box, Dropbox, Google Drive, and OneDrive — don't provide recovery options or data protection (e.g., encryption). A disaster that affects corporate data and transactions is unlikely to spare the administrative tools for these popular storage alternatives, particularly in Software-as-a-Service (SaaS) environments. By default, enterprise IT departments must take up the slack for protecting these and other critical online tools unless their CSP agreements specifically make the CSPs responsible for this component of a cloud system.
As a remedy, IT departments must explore and understand recovery requirements for whichever of these management tools are in use and for any independently purchased productivity platforms. Then, some best-practice procedures for protecting these tools have to be established and made part of the enterprise CDR regimen and documentation. This will likely include storing reliable independent backups of SaaS data and productivity tools. The benefit is knowing that the enterprise cloud environment can be restored entirely on the enterprise's terms and schedule.
While it's impossible to cover all aspects of CDR planning in any cloud environment, hopefully, this overview provides a summary of important basic steps and points out a couple of subtle pitfalls. Avoiding overlooking important details during early planning can save unnecessary expense and provide the strongest assurance against service disruptions from becoming devastating incidents.
Business users want new applications now. Market and regulatory pressures require faster application updates and delivery into production. Your IBM i developers may be approaching retirement, and you see no sure way to fill their positions with experienced developers. In addition, you may be caught between maintaining your existing applications and the uncertainty of moving to something new.
IT managers hoping to find new IBM i talent are discovering that the pool of experienced RPG programmers and operators or administrators with intimate knowledge of the operating system and the applications that run on it is small. This begs the question: How will you manage the platform that supports such a big part of your business? This guide offers strategies and software suggestions to help you plan IT staffing and resources and smooth the transition after your AS/400 talent retires. Read on to learn:
LATEST COMMENTS
MC Press Online