Showing posts with label Advancing Reliability. Show all posts
Showing posts with label Advancing Reliability. Show all posts

Thursday, 25 November 2021

Advancing service resilience in Azure Active Directory with its backup authentication service

The most critical promise of our identity services is ensuring that every user can access the apps and services they need without interruption. We’ve been strengthening this promise to you through a multi-layered approach, leading to our improved promise of 99.99 percent authentication uptime for Azure Active Directory (Azure AD). Today, I am excited to share a deep dive into generally available technology that allows Azure AD to achieve even higher levels of resiliency.

The Azure AD backup authentication service transparently and automatically handles authentications for supported workloads when the primary Azure AD service is unavailable. It adds an additional layer of resilience on top of the multiple levels of redundancy in Azure AD. You can think of it as a backup generator or uninterrupted power supply designed to provide additional fault tolerance while staying completely transparent and automatic to you. This system operates in the Microsoft cloud but on separate and decorrelated systems and network paths from the primary Azure AD system. This means that it can continue to operate in case of service, network, or capacity issues across many Azure AD and dependent Azure services.

What workloads are covered by the service?

This service has been protecting Outlook Web Access and SharePoint Online workloads since 2019. Earlier this year we completed backup support for applications running on desktops and mobile devices, or “native” apps. All Microsoft native apps including Office 365 and Teams, plus non-Microsoft and customer-owned applications running natively on devices are now covered. No special action or configuration changes are required to receive the backup authentication coverage.

Starting at the end of 2021, we will begin rolling out support for more web-based applications. We will be phasing in apps using Open ID Connect, starting with Microsoft web apps like Teams Online and Office 365, followed by customer-owned web apps that use Open ID Connect and Security Assertion Markup Language (SAML).

How does the service work?

When a failure of the Azure AD primary service is detected, the backup authentication service automatically engages, allowing the user’s applications to keep working. As the primary service recovers, authentication requests are re-routed back to the primary Azure AD service. The backup authentication service operates in two modes:

◉ Normal mode: The backup service stores essential authentication data during normal operating conditions. Successful authentication responses from Azure AD to dependent apps generate session-specific data that is securely stored by the backup service for up to three days. The authentication data is specific to a device-user-app-resource combination and represents a snapshot of a successful authentication at a point in time.

◉ Outage mode: Any time an authentication request fails unexpectedly, the Azure AD gateway automatically routes it to the backup service. It then authenticates the request, verifies artifacts presented are valid (such as, refresh token, and session cookie), and looks for a strict session match in the previously stored data. An authentication response, consistent with what the primary Azure AD system would have generated, is then sent to the application. Upon recovery, traffic is dynamically re-routed back to the primary Azure AD service.

Azure Active Directory, Azure Exam Prep, Azure Tutorial and Material, Azure Certification, Azure Career, Azure Guides, Azure Preparation

Routing to the backup service is automatic and its authentication responses are consistent with those usually coming from the primary Azure AD service. This means that the protection kicks in with no need for application modifications, nor manual intervention.

Note that the priority of the backup authentication service is to keep user productivity alive for access to an app or resource where authentication was recently granted. This happens to be most of the type of requests to Azure AD—93 percent, in fact. “New” authentications beyond the three-day storage window, where access was not recently granted on the user’s current device, are not currently supported during outages, but most users access their most important applications daily from a consistent device.

How are security policies and access compliance enforced during an outage?


The backup authentication service continuously monitors security events which affect user access to keep accounts secure, even if these events are detected right before an outage. It uses Continuous Access Evaluation to ensure the sessions that are no longer valid are revoked immediately. Examples of security events that would cause the backup service to restrict access during an outage include changes to device state, account disablement, account deletion, access being revoked by an admin, or detection of a high user risk event. Only once the primary authentication service has been restored would a user with a security event be able to regain access.

In addition, the backup authentication service enforces Conditional Access policies. Policies are re-evaluated by the backup service before granting access during an outage to determine which policies apply and whether the required controls for applicable policies like multi-factor authentication (MFA) have been satisfied. If an authentication request is received by the backup service and a control like MFA has not been satisfied, then that authentication would be blocked.

Azure Active Directory, Azure Exam Prep, Azure Tutorial and Material, Azure Certification, Azure Career, Azure Guides, Azure Preparation
Conditional Access policies that rely on conditions such as user, application, device platform, and IP address are enforced using real-time data as detected by the backup authentication service. However, certain policy conditions (such as sign-in risk and role membership) cannot be evaluated in real-time, and are evaluated based on resilience settings. Resilience defaults enable Azure AD to safely maximize productivity when a condition (such as group membership) is not available in real-time during an outage. The service will evaluate a policy assuming that the condition has not changed since the latest access just before the outage.

While we highly recommend customers to keep resilience defaults enabled, there may be some scenarios where admins would rather block access during an outage when a Conditional Access condition cannot be evaluated in real-time. For these rare cases, administrators can disable resilience defaults per policy within Conditional Access. If resilience defaults are disabled by policy, the backup authentication service will not serve requests that are subject to real-time policy conditions, meaning those users may be blocked by a primary Azure AD outage.

Source: microsoft.com

Saturday, 9 October 2021

Advancing reliability through a resilient cloud supply chain

Microsoft’s cloud supply chain is essential to deliver the infrastructure—servers, storage, and networking gear—that enables cloud reliability and growth. Our vision is for cloud capacity to be available like a utility so that customers can seamlessly turn it on when and where they need it. With the Microsoft Cloud powering everything from mission-critical business applications, governments, life and safety services, financial services and much more, it's crucial that customers are able to scale out when they need it - even with unplanned spikes in demand. To deliver this experience, a resilient, predictable and agile supply chain is key.

Our cloud is growing fast, and we’ll be adding 50 to 100 new datacenters each year for the foreseeable future to the 200 plus datacenters we currently have in operation across 34 countries.   More regions were announced last year than any time previously. The supply chain supports the scale of this growth through an end-to-end value chain of activities. We plan the products, source the materials, build the products, deliver and install them in the datacenters, manage the customer capacity experience, and finally service and decommission the hardware at end of life. Systems and data link the processes so that each activity informs the next.

Azure Cloud, Azure Tutorial and Materials, Azure Learning, Azure Preparation, Azure Career, Azure Prep

How Microsoft is handling supply chain disruptions


The pandemic has disrupted supply chains globally, across many industries. We plan for the unexpected, and yet the pandemic created dramatic changes in demand and supply that we continue to learn and grow from. These changes have required us to respond with agility. Here are some ways the cloud supply chain is building on lessons learned over the past year to strengthen our resiliency and mitigate risks:

◉ End-to-end visibility: Near real-time visibility to supply, inventory and factory status aligned with demand is fundamental to managing supply disruptions and responding to exceptions. We’re investing in this area and launched a “control tower” initiative leveraging Azure and partner services to enable increased visibility and digitally transform our supply chain. Our control tower capability is based on a unified data model built on top of a digital twin of our extended supply chain that will give us finer-grained control over operations. We’ll see more progress in this area over the next year.

◉ Improving lead time for servers: The Azure customer experience is a core priority. In order to maintain the best customer experience while responding to increasing demand, we've improved lead times for building, delivering, and bringing new server capacity live in our datacenters. We have improved these times by 70 percent by automating, accelerating, and removing processes out of the critical path. In addition, we shifted our fulfillment model from build-to-order to build-to-stock for high-volume products. With short lead times, we can respond to customer demand quickly and with agility.

◉ Increased supplier diversification: Over the past year, we have diversified our supply chain—instead of sourcing key subassemblies from one nation, we have diversified our sources to over five countries. This enhances Azure supply chain resiliency to provide stable services to customers during unforeseen circumstances such as the pandemic, natural disasters, or trade challenges.

◉ Intelligent buffering: We have top data scientists and software engineers at Microsoft who have helped us develop an intelligent buffering model, which has allowed us to “shift left” in the supply chain process. This means we carry buffer (inventory) in more agile or raw components rather than finished goods to give us more flexibility to respond to what we need in the datacenters. This has enabled us to better protect our business from demand surges and supply disruptions.
By applying all of these lessons over the last year and a half, we’ve come a long way in strengthening our resiliency to better respond to supply or demand volatility, in turn  improving capacity fulfillment success on our platform and advancing reliability for our customers.

What makes the Microsoft cloud supply chain different?


In addition to our supply chain resilience, Microsoft’s cloud supply chain has unique features that enhance our competitive advantage and our customer experience.

Sustainability

Microsoft has made bold sustainability commitments to be carbon negative, zero waste, and water positive by 2030. The supply chain plays a major role in helping the company achieve these goals. To achieve our zero waste targets, we are building Circular Centers, which are facilities that reuse and repurpose servers and hardware that are being decommissioned. We have one Circular Center in operation, are currently building another, and have nine more on the roadmap. Our Circular Centers use intelligent routing to process and decommission assets on site, maximizing sustainability and value return. We expect the Microsoft Circular Centers to increase the reuse of our servers and components by up to 90 percent by 2025.

To support our carbon-negative goals, we’ve introduced the Sustainability Calculator to enable our customers as well as Microsoft to have transparency into carbon emissions of cloud usage. We’re also serious about carbon accounting to measure and monitor across organizational, process, and product pathways. Every method we create is third-party verified and shared publicly.

Innovation

We seek to be an innovation engine that pushes the supply chain industry forward. We’re investing in our decision science to leverage machine learning, optimization algorithms, artificial intelligence, and digital twins for supply chain so we can make faster, smarter decisions. Our control tower will make it possible to manage all these processes end-to-end.

We’ve also developed a blockchain-based solution for supply chain to improve traceability and create trust in data across our supplier partners by digitizing items in a shared data structure. We won the Gartner Power of the Profession Award for Supply Chain Breakthrough of the Year for our blockchain technology. We’re excited to be currently in production with SSD and DRAM and expanding to the full high-value commodities supplier base. In the future, the technology will enable traceability all the way from mine to datacenter and beyond into recycling, reuse, and disposition.

Trust

We want Azure to be known as the most trusted cloud on the planet. This includes customer trust in Azure as a platform and in Microsoft as a partner in their success. Earning that trust requires two capabilities—security and resilience. We’ve put in place the Azure Security and Resiliency Architecture (ASRA) as an approach to drive security and resiliency consistently and comprehensively throughout the lifecycle of Azure hardware, systems, infrastructure, and supply base.

Microsoft spends $1 billion a year on security, and our hardware and datacenters are designed with security top of mind. Microsoft is leading the way in confidential computing, deploying hardware that is physically and logically isolated from someone that has access to the server while it’s operating.
Finally, we are building transparency into our supply chain so that we know we are sourcing ethically and are accountable partners along each step of the way. This includes risk management, anti-bribery and anti-corruption, human and labor rights, health and safety, and more.

Talent

Our strength is in our people, and we want to be known as a talent destination. At Microsoft we expect each of us to play a role in creating an inclusive environment where people of diverse backgrounds bring all of who they are to our work. Different perspectives help us achieve more and inclusive thinking drives our innovation.

Azure Cloud, Azure Tutorial and Materials, Azure Learning, Azure Preparation, Azure Career, Azure Prep

We’re also investing in our people and capabilities through a new initiative called Supply Chain Academy, which develops our muscle around supply chain excellence by offering online courses on supply chain best practices and disciplines.

It’s all about our customers


At the end of the day, everything we do comes back to providing the best experience for our customers. By investing in our agility, resiliency, innovation, security, and talent, we are building a world-class supply chain that will make Azure the most reliable and trusted cloud platform. We’re grateful to all our partners, suppliers, and customers who have joined us in this journey as we power the world’s computers and empower every person and organization on the planet to achieve more.

Tuesday, 3 August 2021

Advancing Azure Virtual Machine availability transparency

The existing Azure resource health feature helps you to diagnose and get support for service problems that affect your Azure resources. It reports on the current and past health of your resources, showing any time ranges that each of your resources have been unavailable. But we know that our customers and partners are particularly interested in “the why” to understand what caused the underlying technical issue, and in improving how they can receive communications about any issues—to feed into monitoring processes, to explain hiccups to other stakeholders, and ultimately to inform business decisions.

Read More: MS-500: Microsoft 365 Security Administration

Introducing root causes for VM issues—in Azure resource health

We recently shipped an improvement to the resource health experience that will enhance the information we share with customers about VM failures, with additional context on the root cause that led to the issue. Now, in addition to getting a fast notification when a VM’s availability is impacted, customers can expect a root cause to be added at a later point once our automated Root Cause Analysis (RCA) system identifies the failing Azure platform component that led to the VM failure. Let’s walk through an example to see how this works in practice:

1. At time T1, a server rack goes offline due to a networking issue, causing VMs on the rack to lose connectivity. (Recent reliability improvements related to network architecture will be shared in a future Advancing Reliability blog post—watch this space!)

2. At time T2, Azure’s internal monitoring recognizes that it is unable to reach VMs on the rack and begins to mitigate by redeploying the impacted VMs to a new rack. During this time, an annotation is sent to resource health notifying customers that their VM is currently impacted and unavailable.

Azure Virtual Machine, Azur Exam Prep, Azure Tutorial and Material, Azure Certification, Azure Preparation, Azure Career
Figure 1: A screenshot of the Azure portal “resource health” blade showing the health history of a resource.

3. At time T3, platform telemetry from the top of rack switch, the host machine, and internal monitoring systems, are all correlated together in our RCA engine to derive the root cause of the failure. Once computed, the RCA is then published back into resource health along with relevant architectural resiliency recommendations that customers can implement to minimize the probability of impact in the future.

Azure Virtual Machine, Azur Exam Prep, Azure Tutorial and Material, Azure Certification, Azure Preparation, Azure Career

Figure 2: A screenshot of the Azure portal "health history" blade showing root cause details for an example of a VM issue.

While the initial downtime notification functionality has existed for several years, the publishing of a root cause statement is a new addition. Now, let’s dive into the details of how we derive these root causes.

Root Cause Analysis engine


Let’s take a closer look at the prior example and walk through the details of how the RCA engine works and the technology behind it. At the core of our RCA engine for VMs is Azure Data Explorer (ADX), a big data service optimized for high volume log telemetry analytics. Azure Data Explorer enables the ability to easily parse through terabytes of log telemetry from devices and services that comprise the Azure platform, join them together, and interpret the correlated information streams to derive a root cause for different failure scenarios. This ends up being a multistep data engineering process:

Phase 1: Detecting downtime

The first phase in root cause analysis is to define the trigger under which the analysis is executed. In the case of Virtual Machines, we want to determine root causes whenever a VM unexpectedly reboots, so the trigger is a VM transitioning from an up state to a down state. Identifying these transitions from platform telemetry is straightforward in most scenarios, but more complicated around certain kinds of infrastructure failure where platform telemetry might get lost due to device failure or power loss. To handle these classes of failures, other techniques are required—like tracking data loss as a possible indication of a VM health transition. Azure Data Explorer excels at this time of series analysis, and a more detailed look at techniques around this can be found in the Microsoft Tech Community: Calculating downtime using Window functions and Time Series functions in Azure Data Explorer.

Phase 2: Correlation analysis

Once a trigger event is defined (in this case, a VM transitioning to an unhealthy state) the next phase is correlation analysis. In this step we use the presence of the trigger event to correlate telemetry from points across the Azure platform, like:

◉ Azure host: the physical blade hosting VMs.
◉ TOR: the top of rack network switch.
◉ Azure Storage: the service which hosts Virtual Disks for Azure Virtual Machines.

Each of these systems has their own telemetry feeds that need to get parsed and correlated with the VM downtime trigger event. This is done through understanding the dependency graph for a VM and the underlying systems that can cause a VM to fail, and then joining all these dependent systems’ health telemetry together, filtered on events that are relatively close to the VM transition in time. Azure Data Explorer’s intuitive and powerful query language helps with this, with documented patterns like time window join for correlating temporal telemetry streams together. At the end of this correlation process, we have a dataset that represents VM downtime transitions with correlated platform telemetry from all the dependent systems that could cause or could have information useful in determining what led to the VM failure.

Phase 3: Root cause attribution

The next step in the process is attribution. Now that we’ve collected all the relevant data together in a single dataset, attribution rules get applied to interpret the information and translate it into a customer-facing root cause statement. Going back to our original example of a TOR failure, after our correlation analysis we might have many interesting pieces of information to interpret. For example, systems monitoring the Azure hosts might have logs indicating they lost connectivity to the hosts during this time. We might also have signals related to virtual disk connectivity problems, and explicit signals from the TOR device about the failure. All these pieces of information are now scanned over, and the explicit TOR failure signal is prioritized over the other signals as the root cause. This prioritization process, and the rules behind it, are constructed with domain experts and modified as the Azure platform evolves. Machine learning and anomaly detection mechanisms sit on top of these attributed root causes, to help identify opportunities to improve these classification rules as well as to detect pattern changes in the rate of these failures to feed back into safe deployment pipelines.

Phase 4: RCA publishing

The last step is publishing root causes to Azure resource health, where they become visible to customers. This is done in a very simple Azure Functions application, which periodically queries the processed root cause data in Azure Data Explorer, and emits the results to the resource health backend. Because information streams can come in with various data delays, RCAs can occasionally be updated in this process to reflect better sources of information having arrived leading to a more specific root cause that what was originally published.

Going forward


Identifying and communicating to our customers and partners the root cause of any issues impacting them, is just the beginning. Our customers may need to take these RCAs and share them with their customers and coworkers. We want to build on the work here to make it easier to identify and track resource RCAs, as well as easily share them out. In order to accomplish that, we are working on backend changes to generate unique per-resource and per-downtime tracking IDs that we can expose to you, so that you can easily match downtimes to their RCAs. We are also working on new features to make it easier to email RCAs out, and eventually subscribe to RCAs for your VMs. This will make it possible to sign up for RCAs directly in your inbox after an unavailability event with no additional action needed on your part.

Source: microsoft.com

Saturday, 24 July 2021

Advancing global network reliability through intelligent software - Part 2

Read Advancing global network reliability through intelligent software part 1.

In part one of this networking post, we presented the key design principles of our global network, explored how we emulate changes, our zero touch operations and change automation, and capacity planning. In part two, we start with traffic management. For a large-scale network like ours, it would not be efficient to use traditional hardware managed traffic routing. Instead, we have developed several software-based solutions to intelligently manage traffic engineering in our global network.

SDN-based Internet Traffic Engineering (ITE)

The edge is the most dynamic part of the global network—because the edge is how users connect to Microsoft’s services. We have strategically deployed edge PoPs close to users, to improve customer/network latency and extend the reach of Microsoft cloud services.

Read More: MS-900: Microsoft 365 Fundamentals

For example, if a user in Sydney, Australia accesses Azure resources hosted in Chicago, USA their traffic enters the Microsoft network at an edge PoP in Sydney then travels on our network to the service hosted in Chicago. The return traffic from Azure in Chicago flows back to Sydney on our Microsoft network. By accepting and delivering the traffic to the point closest to the user, we can better control the performance.

Azure Exam Prep, Azure Tutorial and Materials, Azure Preparation, Azure Learning, Azure Guides, Azure Career

Each edge PoP is connected to tens or hundreds of peering networks. Routes between our network and providers’ networks are exchanged using the Border Gateway Protocol (BGP). BGP best path selection has no inherent concept of congestion or performance—neither is BGP capacity aware. So, we developed an SDN-based Internet Traffic Engineering (ITE) system that steers traffic at the edge. The entry and exit points are dynamically altered based on the traffic load of the edge, internet partners’ capacity constraints, reduction or augments in capacity, demand spikes sometimes caused by distributed denial of service attacks, and latency performance of our internet partners. The ITE controller constantly monitors these signals and alters the routes we advertise to our internet partners and/or the routes advertised inside the Microsoft network, to select the best peer-edge.

Optimizing last mile resilience with Azure Peering Service


In addition to optimizing routes within our global network, the Azure Peering Service extends the optimized connectivity to the last mile in the networks of Internet Service Providers (ISPs). Azure Peering Service is a collaboration platform with providers, to enable reliable high-performing connectivity from the users to the Microsoft network. The partnership ensures local and geo redundancy, and proximity to the end users. Each peering location is provisioned with redundant and diverse peering links. Also, providers interconnect at multiple Microsoft PoP locations so that if one of the edge nodes has degraded performance, the traffic routes to and from Microsoft via alternative sites. Internet performance telemetries from Map of Internet (MOI) drive traffic steering for optimized last mile performance.

Route Anomaly Detection and Remediation (RADAR)


The internet runs on BGP. A network or autonomous system is bound to trust, accept, and propagate the routes advertised by its peers without questioning its provenance. That is the strength of BGP and allows the internet to update quickly and heal failures. But it is also its weakness—the path to prefixes owned by a network can be changed by accident or malicious intent to redirect, intercept, or blackhole traffic. There are several incidents that happen to every major provider and some make front page news. We developed a global Route Anomaly Detection and Remediation (RADAR) system to protect our global network.

RADAR detects and mitigates Microsoft route hijacks on the Internet. BGP route leak is the propagation of routing announcement(s) beyond their intended scope. RADAR detects route leaks in Azure and the internet. It can identify stable versus unstable versions of a route and validate new announcements. Using RADAR, and the ITE controller, we built real-time protection for Microsoft prefixes. Peering Service platform extends the route monitoring and protection against hijacks, leaks and any other BGP misconfiguration (intended or not) in the last mile up to the customer location.

Software-driven Wide Area Network (SWAN)


The backbone of the global network is analogous to a highway system connecting major cities. The SWAN controller is effectively the navigation system that assigns the routes for each vehicle, such that every vehicle reaches its destination as soon as possible and without causing congestion on the highways. The system consists of topology discovery, demand prediction, path computation, optimization, and route programming.

Azure Exam Prep, Azure Tutorial and Materials, Azure Preparation, Azure Learning, Azure Guides, Azure Career

Over the last 12 months, the speed of the controller to program the network improved by an order of magnitude and the route-finding capability improved two-fold. Link failures are like lane closures so the controller must recompute routes to decrease congestion. The controller uses the same shared risk link groups (SRLGs) to compute backup routes in case of failure of the primary routes. The backup routes activate immediately upon failure, and the controller gets to work at reoptimizing traffic placement. Links that go up and down in rapid succession are held back from service until they stabilize.

One measure of reliability is the percentage of successfully transmitted bytes to requested bytes, measured over an hour and averaged for the day. Ours is 99.999 percent or better for customer workloads. All communication between Microsoft services is through our dedicated global network. Thousand Eyes Cloud Performance Benchmark reports that over 99 percent of Azure inter-region latencies faster than the performance baseline, and over 60 percent of region pairs are at least 10 percent faster. This is a result of the capacity augments and software systems described in this post.

Bandwidth Broker—software-driven Network Quality of Service (QoS)


If the global network is a highway system, Bandwidth Broker is the system that controls the metering lights at the onramps of highways. For every customer vehicle, there is more than one Microsoft vehicle traversing the highway. Some of the Microsoft vehicles are discretionary and can be deferred to avoid congestion for customer vehicles. Customer vehicles always have a free pass to enter the highways. The metering lights are green in normal operation but when there is a failure or a demand spike, Bandwidth Broker turns on the metering lights in a controlled manner. Microsoft internal workloads are divided into traffic tiers, each with a different priority. Higher priority workloads are admitted in preference to lower priority workloads.

Brokering occurs at the sending host. Hosts periodically request bandwidth on behalf of applications running on them. The requests are aggregated by the controller, bandwidth is reserved, and grants are disseminated to each host. Bandwidth Broker and SWAN coordinate to adjust traffic volume to match routes, and traffic routes to match volume.

It is possible to experience multiple fiber cuts or failures that suddenly reduce network capacity. Geo-replication operations to increase resilience can cause a huge surge in network traffic. Bandwidth Broker generally allows us to preserve the customer experience during these conditions, by shedding discretionary internal workloads when congestion was imminent.

Continuous monitoring


A robust monitoring solution is the foundation to achieve higher network reliability. It lowers both the time to detect and time to repair. The monitoring pipelines constantly analyze several telemetry streams including traffic statistics, health signals, logs, and device configurations. The pipelines automatically collect more data when anomalies are detected or diagnose and remediate common failures. These automated interventions are also guarded by safety check systems.

Major investments in monitoring have been:

➤ Polling and ingestion of metrics data at sub-minute speeds. A few samples are needed to filter transients and a few more to generate a strong signal. This leads to faster detection times.
➤ An enhanced diagnostics system that is triggered by packet loss or latency alerts, instructs agents at different vantage points to collect additional information to help triangulate and pinpoint the issue to a specific link or device.
➤ Enhanced diagnostics trigger auto-mitigation and remediation actions for the most common incidents, with the help of Clockwerk and Real Time Operation Checker (ROC). This translates to faster time to repair and has the ripple effect of keeping engineers focused on more complex incidents.

Other pipelines continuously monitor network graphs for node isolation, and periodically assess risks with “what-if” intent using ROC as described above. We have multiple canary agents deployed throughout the network checking reachability, latency, and packet loss across our regions. This includes agents within Azure, as well as outside of our network, to enable outside-in monitoring. We also periodically analyze Map of Internet (MOI) telemetries to measure end to end performance from customers to Azure. Finally, we have robust monitoring in place to protect the network from security attacks such as BGP route hijacks, and distributed denial of service (DDoS).

Source: Microsoft.com

Tuesday, 13 July 2021

Advancing application reliability with the Azure Well-Architected Framework

If you want to start a good discussion or argument about reliability at work, ask a colleague this question.

"When is architecture more important for the reliability of a service, product, or application? Before it is deployed to production, or afterward?"

Well, “surely”—you say—“if we don’t build the service with reliability in mind, it may not have the right components included to increase stability. It may not have redundancy to improve fault tolerance. Perhaps we will have left out robust retry logic, circuit breakers, or other known patterns for reliable systems.”

But maybe your colleague counters, “Well, I can’t deny that it is important to attempt to try and build things right from the beginning. But one thing I’ve learned about reliability is it is almost never achieved on the first go around. Even if you have done a phenomenal job at the whiteboard, designing with failure in mind, there are still going to be outages. And while nobody likes outages, if we handle them and a subsequent post-incident review correctly, we can learn a great deal that helps us make a service more reliable in the long term. On top of this, wouldn’t you agree that observability is an iterative process that involves changing what we measure and monitor as we learn more about the system while it is running? All these things would fall under the mantle John Reese and Niall Murphy called 'the wisdom of Production'. And all of these things surely need us to bring to bear all the architecture skills we have to do this right.”

If you are having a really good discussion, this goes back and forth across the table at least a few times. One side notes that “bolting on reliability after the fact” works about as well as “bolting on security after the fact” (that is to say, not well at all). The other side might bring up the lessons we’ve learned from chaos engineering showing us that experiments on a dev or staging environment can be very useful, but they don’t always yield some of the unique results we get from testing in production.

“But what about the value of continuous integration and continuous delivery (CI/CD) to reliability—trying to catch reliability issues before they get to production?”, gets asked. Then in response, “CI/CD is tremendously useful, but it didn’t catch our last issue because tests for large distributed systems are notoriously hard to get right.” And so on, and so on.

By now you’ve probably come to the same conclusion the people in this argument are bound to reach. Architecture is important in both the pre-production and post-production lifecycle stages. But that conclusion still leaves us in a peculiar spot because we don’t normally think about architecture or the role of an architect after something has been built. We don’t expect the architect who helped us build our house to show up at the doorstep a year later to say “OK, let’s do some more architecting.”

With the applications we build (or purchase) to run, things are different. There we have an expectation that the software will be changed at a much more rapid pace. It will be refactored, it will be enhanced, it will be upgraded. At each of these points, we must apply everything we know from the realm of architecture if we expect the result to be reliable. So let me tell you about one way to settle the debate we’ve been discussing, and also show you a tool that can help with your reliability even as we are squaring that circle.

The Azure Well-Architected Framework

The Well-Architected Framework is a set of guiding tenets that can be used to improve the quality of a workload. The framework consists of five pillars of architecture excellence: Cost Optimization, Operational Excellence, Performance Efficiency, Reliability, and Security. Incorporating these pillars helps produce high-quality, stable, and efficient cloud architecture.

But there’s that word “architecture” again, basically sitting right in the middle of the name and taunting us with an image of an architect who only participates at the beginning of the lifecycle.

Here’s the key to unlocking this conundrum: For reliability (and the other four pillars) the goal is to work towards and remain in a “well-architected state.”

That’s a state that strives to embody and make use of the best practices and all the accumulated knowledge from architecture meticulously embedded in the Well-Architected Framework. This guidance is meant to be useful to you at all stages of a cloud solution. It is useful to you in the beginning when you are designing your workloads. It is useful to you when you begin your periodic review of the workload as part of the refactoring, scaling, enhancing, or upgrading process. And finally, it can help when the cycle starts anew for the next major version of your workload.

How to get there

Anyone who has worked in the reliability space, even for a short while, knows that while a large body of guidance like the Well-Architected Framework is great, the tricky part is applying that knowledge to your specific workloads and efforts in flight. Just navigating a large document set like the Well-Architected Framework and determining where to start can be a challenge. I’d like to introduce you to a tool that I believe can bridge your ground truth and the guidance we offer. It can serve as our compass to this material.

Azure Well-Architected Framework, Azure Exam Prep, Azure Tutorial and Materials, Azure Learning, Azure Prep, Azure Preparation, Azure Guides

The Well-Architected Review is a self-guided assessment tool that will walk you through the Well-Architected Framework reliability pillar and the other four Well-Architected Framework pillars. This is a great process to do either by yourself, with your friendly neighborhood Cloud Solutions Architect, or supporting partners. It will ask you a set of questions about your reliability efforts—then, based on your responses, it offers suggestions on areas to focus on with direct links to our WAF documentation on those areas.

Here’s an example set of results:

Azure Well-Architected Framework, Azure Exam Prep, Azure Tutorial and Materials, Azure Learning, Azure Prep, Azure Preparation, Azure Guides

Let me offer a few tips that might not be obvious at first look for getting the most out of the Well-Architected Review:

1. Pay attention to the questions: You might think the results of the review are the biggest reward, but I’m here to tell you that the most valuable thing you may be able to take away from the review is the questions. Reliability can be a tricky area to tackle because there are so many possible ways to begin working on it, so many different places to start. Just knowing which questions to ask can be difficult. The Well-Architected Review can give you those questions.

2. Return to the review again and again: If you sign into the review platform with your Microsoft credentials, you can save the results. This means that in six months, or whenever you feel ready to conduct another review, you will be able to compare your new review to your previous information. This can be tremendously helpful for judging your progress across each pillar.

3. Share the results with your team: One thing many people don’t know about the Well-Architected Review is if you have signed in (see tip number two above), it will allow you to export your results as a Microsoft PowerPoint presentation. Take this draft, customize it, and you now have a ready-made presentation to take to your next team meeting so everyone can get behind your reliability efforts.

The Well-Architected Framework in action

If you would like to see some examples of the Well-Architected Framework in action, including some excellent sessions about reliability, I encourage you to check out the videos in the Well-Architected series of our Azure Enablement show. There’s some good online course material about the subject in Microsoft Learn, and guiding principles in our documentation. If you want to dive deeper into the architecture side of Well-Architected, I recommend checking out the Azure Architecture Center.

Source: microsoft.com

Saturday, 10 July 2021

Advancing resiliency threat modeling for large distributed systems

Azure Exam Prep, Azure Learning, Azure Tutorial and Material, Azure Certification, Azure Study Material, Azure Prep

At a high level, our goal is to avoid self-inflicted and/or avoidable outages, but even more immediately, our goal is to be able to reduce the likelihood of impact to our customers as much as possible. In this context, an outage is any incident in which our services fail to meet the expectations of (or negatively impact the workloads of) our customers, both internally and externally. To avoid that, we needed to improve how we uncover risks before they impact customer workloads on Azure. But Azure itself is a large, complex distributed system—how should we approach resiliency threat modeling if our organization offers thousands of solutions, comprised of hundreds of “services,” each with a team of five to 50 engineers, distributed across multiple different parts of the organization, and each with their own processes, tools, priorities, and goals? How do we scale our resiliency threat modeling process out, and reason across all these individual risk assessments? To address these challenges, it took some major changes to join the reactive approach with the more proactive approach.

Starting our journey

We started the shift left with a premortem pilot program. We looked back at past outages and developed a questionnaire that helped not only to start discussions but also to provide a structure to them. Next, we selected several services of varying purpose and architecture—each time we sat down with a team, we learned something new, got better and better at identifying risks, incorporated feedback from the teams on the process, then tried again with the next team. Eventually, we started to identify the right questions to ask as well as other elements we needed to make this process productive and impactful. Some of these elements already existed, others needed to be created to support a centralized approach to resiliency threat modeling. Many that existed also required changes or integration into an overall solution that met our goals. What follows is a high-level overview of our approach, and the elements we discovered were necessary to continue improving the space.

Read More: MS-203: Microsoft 365 Messaging

We created a culture of continuous and fearless risk documentation

We realized that there would need to be a large up-front investment to find out where our risks were. Thorough premortems take time—the valuable and "in high demand" engineer kind of time, the "we are working within a tight timeline to deliver customer value and can’t stop to do this" kind of engineer time. Sometimes we needed to help them understand that, even though their service only had one outage in two years, there are dozens of other services in the dependency chains of our solutions that have also only had one outage in the past two years. The point is that, from our customers’ perspective, they have seen more than “just” two outages.

We should be skeptical that low risk is something we can safely ignore. We must embark on healthy, fearless searching inventory of risks to outages. We needed a process that not only supports detailed discussions of lingering issues, and an assessment of the risks without fear of reprisal, but also informs investments in common solutions and mitigations that can be leveraged broadly across multiple services in a large-scale organization consisting of hundreds or thousands of services.

Our goal is a continuous search for risks, and to have "living risks" threat models instead of static models that are only updated every X months. Once the initial investment to find that initial set of risks lands, we wanted to keep things up to date with a well-defined process. The key at first was not to be deterred by how large the initial investment "moat" was, between us and our goals. The real benefit we see now is that risk collection is built into our culture. This enables risks to be collected from many diverse sources and built into our organizational processes.

So, how best to start the process of identifying risks?

We used both reactive and proactive approaches to uncovering risks

We leveraged our postmortem analysis program

While looking for risks, it makes sense to look at what has happened in the past and search for clues to what can happen again in the future. A solid postmortem analysis program was key in this regard. We already had a team that analyzed postmortems, looking for common themes and surfacing them for deeper analysis and to inform investments. Their analyses helped us not only route teams to quality programs or initiatives that could help address the risk, but also highlighted the need for other investments we did not know we needed until we saw how prevalent the risk was. The types of risk categories they identified from postmortem analysis became a subset of our focus areas when we looked for risks that had not happened yet. It is worth noting here that postmortem analysis is only as good as the postmortem itself.

We leveraged our postmortem quality review program

The quality of a postmortem determines its usefulness, therefore we invested in a Postmortem Quality Review Program. Guidance was published, training was made available, and a large pool of reviewers would rate each "high impact" outage postmortem after it was written. Postmortems that had low ratings or needed more clarity were sent back to the authors. High impact postmortems are reviewed weekly in a meeting that includes engineers from other teams and senior leaders, both of whom ask questions and give feedback around the right action plans. Having this program vastly increased our ability to learn from and act on postmortem data.

However, we knew that we could not limit our search for risks to the past.

We got better at premortems to be more proactive

Some may have heard the term “premortem,” which is a similar process to a Failure Mode Analysis (FMA).

According to the Harvard Business Review (September 2007, Issue 1), “A premortem is the hypothetical opposite of a postmortem. A postmortem in a medical setting allows health professionals and the family to learn what caused a patient’s death. Everyone benefits except, of course, the patient. A premortem in a business setting comes at the beginning of a project rather than the end so that the project can be improved rather than autopsied. Unlike a typical critiquing session, in which project team members are asked what might go wrong, the premortem operates on the assumption that the “patient” has died, and so asks what did go wrong. The team members’ task is to generate plausible reasons for the project’s failure.”

In the context of Azure, the goal of a premortem is to predict what could go wrong, identify the potential risks, and mitigate or remove them before the trigger event results in an outage. According to How to Catch a Black Swan: Measuring the Benefits of the Premortem Technique for Risk Identification, premortem techniques can identify more quality risks and propose more quality changes; it is a more effective risk management technique compared with others.

Conducting a premortem is not a terribly difficult endeavor, but you must be sure you involve the right set of people. For a particular customer solution, we gathered the most knowledgeable engineers for that solution and, even better, brought the new ones so they can learn. Next, we had them brainstorm as many reasons as possible why they may be woken up at 4:00 AM because their customers are experiencing an outage that they must mitigate. We like to start with a hypothetical question, “Imagine you were away on vacation for two weeks and came back hearing that your service had an outage. Before you find out what the contributing factors were for that outage, attempt to list as many things as possible you think could be the likely causes.” Each of those was captured as a risk, including documenting the triggers that cause the outage, the impact on customers, and the likelihood of it happening.

After we identified the risks, we needed to have an action plan to address them.

We created “Risk Threat Models”

If premortem is the searching and fearless inventory of risks, the risk threat model is what combines that with the action plan. If your team is doing failure mode analysis or regular risk reviews, this effort should be straightforward. The risk threat model takes your list of risks and builds in what you expect to do to reduce customer pain.

After we identified the risks, we needed to have a common understanding of the right fixes so that those patterns could be used everywhere there was a similar risk. If those fixes are long-term, we asked ourselves “what can we do in the meantime to reduce the likelihood or impact to our customers”? Here are some of the questions we asked:

◉ What telemetry exists to recognize the risk when it is triggered? 

◉ What if this happens tomorrow? What guardrails or processes are in place until the final fix is completed?

◉ How long does it take to do this interim mitigation and is it automated? Can it be automated? Why not?

◉ Has the mitigation been tested? When was the last time you tested it?

◉ What have you done to ensure you have added the most resiliency possible, for instance, if your critical dependencies have an outage? Have you worked with that dependency to ensure you are following their guidance and using their latest libraries?

◉ Are there any places where you can ask for an increase in COGS to be more resilient?

◉ What mitigations are you not doing due to the cost or complexity or because you do not have the resources?

Risk Mitigation Plans were documented, aligned with common solutions, and tracked. Any teams that indicated the mitigations would need more developers or money to implement needed mitigations were provided a forum to ask for it in the context of resiliency and quality. Work items were created for all the tasks needed to make the service as resilient as possible and linked to the risks we documented so we could follow up on their completion.

But how do we know we are precluding the risk correctly? How do we know what the "right" mitigation strategies are? How could we reduce the amount of work it took, and prevent everyone from solving the same problem in a custom way?

We created a centralized repository of all risks across all services

Risks Threat Models are more effective when captured and analyzed centrally

If every team went off and analyzed their postmortems, conducted a premortem, created a risk threat model, and then documented everything separately there is no doubt we would be able to reduce the number of outages we have. However, by not having a way for teams to share what they found with others, we would have missed opportunities to understand broader patterns in the risk data. Risks in one service were often risks in other services as well, but for one reason or another, it did not make it into the risk threat models of all services that were at risk. We also would not have noticed that unfinished repairs or risks in one service are risks to other services. We would have missed the chance to document patterns that will prevent this risk from appearing elsewhere when a new service is spun up. We would not have realized that the scope of a premortem should not be limited to individual service boundaries, but rather done across all the services that work together for a particular customer scenario. We would have missed opportunities to propose common mitigation strategies and invest in broad efforts to address mitigations at scale using the same mitigation plan. In short, we would have missed many opportunities to inform numerous investments across service boundaries.

We implemented a common way of categorizing risks

In some cases, we knew we needed to understand what "types" of risks we were finding. Were they Single Points of Failure? Insufficiently configured throttling? To that end, we created a hierarchical system of shorthand “tags” that were used to describe categories of issues on which we wanted to focus. These tags were used for analyzing postmortems to identify common patterns, as well as marking individual risks so that we could better look across risks in the same categories to identify the right action plans.

We had regular reviews of the Risk Threat Models

Having the completed Risk Threat Models enabled us to schedule reviews in front of senior leadership, architects, members of the dedicated cross-Azure Quality Team, and others. These meetings were more than just reviews, they provided an opportunity to come together as a diverse team to identify areas for which we needed common solutions, mitigations, and follow-up actions. Action items were collected, owners assigned, and risks were then linked with the action plans so we could follow up down the road to determine how teams progressed.

It all came together, time to take it to the next level and do this for hundreds of services!

In summary, it took more than just spinning up a program to identify and document risks. We needed to inspire, but also have the right processes in place to get the most out of that effort. It took coordination across many programs and the creation of many others. It took a lot of cross-service-team communication and commitment.

Accelerating the resiliency threat modeling Program has already yielded many benefits for our critical Azure services, so we will be expanding this process to cover every service in Azure. To this end, we are continuously refining our process, documentation, and guidance as well as leveraging past risk discussions to address new risks. Yes, this is a lot of work, and there is no silver bullet, and we are still bringing more and more resources into this effort, but when it comes to reliability, we believe in “go big”! 

Source: microsoft.com

Thursday, 1 July 2021

Advancing safe deployment with AIOps—introducing Gandalf

In our earlier blog post “Advancing safe deployment practices” Cristina del Amo Casado described how we release changes to production, for both code and configuration changes, across the Azure platform. The processes consist of delivering changes progressively, with phases that incorporate enough bake time to allow detection at a small scale for most regressions missed during testing.

Read More: PL-900: Microsoft Power Platform Fundamentals

The continuous monitoring of health metrics is a fundamental part of this process, and this is where AIOps plays a critical role—it allows the detection of anomalies to trigger alerts and the automation of correcting actions such as stopping the deployment or initiating rollbacks.

In the post that follows, we introduce how AI and machine learning are used to empower DevOps engineers, monitor the Azure deployment process at scale, detect issues early, and make rollout or rollback decisions based on impact scope and severity.

Why AIOps for safe deployment

As defined by Gartner, AIOps enhances IT operations through insights that combine big data, machine learning, and visualization to automate IT operations processes, including event correlation, anomaly detection, and causality determination. In our earlier post, "Advancing Azure service quality with artificial intelligence: AIOps," we shared our vision and some of the ways in which we are already using AIOps in practice, including around safe deployment. AIOps is well suited to catching failures during deployment rollout, particularly because of the complexities of cross-service dependencies, the scale of hyperscale cloud services, and the variety of different customer scenarios supported.

Phased rollouts and enriched health signals are used to facilitate monitoring and decision making in the deployment process, but the volume of signals and level of complexity involved in deployment decision making exceeds what any human could reasonably reason over, across thousands of ever-evolving service components, spanning more than 200 datacenters in more than 60 regions. Some latent issues won’t manifest for several days after their deployment, and global issues that span different clusters but manifest only minutely in any individual cluster are hard to detect with just a local watchdog. While loose coupling allows most service components to be deployed independently, their deployments could have intricate impacts on each other. For example, a simple change in an upstream service could potentially impact a downstream service if it breaks the contract of API calls between the two services.

These challenges call for automated monitoring, anomaly detection, and rollout impact assessment solutions to facilitate deployment decisions at velocity.

Figure 1: Gandalf safe deployment

Gandalf safe deployment service: An AIOps solution


Rising to the challenge described above, the Azure Compute Insights team developed the “Gandalf” safe deployment service—an end-to-end, continuous monitoring system for safe deployment. We consider this part of the Gandalf AIOps solution suite, which includes a few other intelligent monitoring services. The code name Gandalf was inspired by the protagonist from The Lord of the Rings, as shown in Figure 1, it serves as a global watchdog, which makes intelligent deployment decisions based on signals collected. It works in tandem with local watchdogs, safe deployment policies, and pre-qualification tests, all to ensure deployment safety and velocity.

As illustrated in Figure 2, the Gandalf system monitors rich and representative signals from Azure, performs anomaly detection and correlation, then derives insights to support deployment decision making and actions.

Figure 2: Gandalf system overview

Data sources

Gandalf monitors signals across performance, failures, and events as described below. It pre-processes the data to structure them around a unified data schema to support downstream data analytics. It also leverages a few other analytics services within Azure for health signals, including our Virtual Machine failure categorization service and near real-time failure attribution processing service. Signal registration with Gandalf is required when any new service components are onboarded, to ensure complete coverage.

◉ Performance data: Gandalf monitors performance counters, CPU usage, memory usage, and more – all for a high-level view of performance and resource consumption patterns of hosted services.

◉ Failure signals: Gandalf monitors both the hosting environment of customer’s virtual machines (data plane) and tenant-level services (control plane). For the data plane, it monitors failure signals such as OS crashes, node faults, and reboots to evaluate the health of the VM’s hosting environment.  At the same time, it monitors failure signals of the control plane like API call failures, to evaluate the health of tenant-level services.

◉ Update events: In addition to telemetry data collected, Gandalf also keeps its finger on the pulse of deployment events, which report deployment progress and issues.

Detection, correlation, and decision

Gandalf evaluates the impact scope of the deployment—for example, the number of impacted nodes, clusters, and customers—to make a go/no-go decision using decision criteria that are trained dynamically. To balance speed and coverage, Gandalf utilizes an architecture with both streaming and batch analysis engines.

Figure 3: Gandalf Anomaly Detection and Correlation Mode

Figure 3 shows an overview of the Gandalf Machine Learning (ML) model. It consists of two parts—anomaly detection and correlation process (to identify suspicious deployments) and a decision process (to evaluate customer impact).

Anomaly detection and correlation process

To ensure precise detection, Gandalf derives fault signatures from input signals, which can be used to uniquely identify the failure. Then, it detects based on the occurrence of the fault signature.

In large-scale cloud systems like Azure, simple threshold-based detection is not practical both because of the dynamic nature of the systems and workloads hosted and because of the sheer volume of fault signatures. Gandalf applies machine learning techniques to estimate baseline settings based on historical data automatically and can adapt the setting through training as needed.

When Gandalf detects an anomaly, it correlates the observed failure with deployment events and evaluates its impact scope. This helps to filter out failures caused by non-deployment reasons such as random firmware issues.

Since multiple system components are often deployed concurrently, a vote-veto mechanism is used to establish the relationship between the faults and the rollout components. In addition, temporal and spatial correlations are used to identify the components at fault. Fault age, which measures the time between rollout and detection of fault signature, is considered to allow more focus on new rollouts than old ones since newly observed faults are less likely to be triggered by the old rollout.

In this way, Gandalf can detect an anomaly that would lead to potential regressions in the customer experience early in the process—before it generates widespread customer impact. 

Decision process

Finally, Gandalf evaluates the impact scope of the deployment such as the number of impacted clusters/nodes/customers, and ultimately makes a "go/no-go" decision. It’s worth mentioning that Gandalf is designed to allow developers to customize signals’ weight assignment based on their experience. In this way, it can incorporate domain knowledge from human experts to complement its machine learning solutions.

Result orchestration

To balance speed and coverage, Gandalf utilizes both streaming and batch processing of incoming signals. Streaming processing consumes data from Azure Data Explorer, a cloud storage solution supporting analytics with fast speed. Streaming processing is used to process fault signals that happen 1 hour before and after each deployment in each node and runs lightweight analysis algorithms for rapid response.

Batch processing consumes data from Cosmos, a Hadoop-like file system that supports extremely large volumes of data. It’s used to analyze faults over a larger time window (generally a 30-day period) with advanced algorithms.

Both stream and batch processing are performed incrementally with five-minute intervals. In general, the incoming telemetry signals of Gandalf are both streamed into Kusto and stored into Cosmos hourly/daily. With the same data source, occasionally there could be inconsistent results from the processing pipeline. This is by design since batch processing makes more informed decisions and covers latent issues that the fast/streaming process cannot detect.

Deployment experience transformation


The Gandalf system is now well integrated into our DevOps workflow within Azure and has been widely adopted for deployment health monitoring across the entire fleet. It not only helps to prevent bad rollouts as quickly as possible but has also transformed the engineers’ and release managers’ experience in deploying software changes—from looking for scattered evidence to using a single source of truth, from ad-hoc diagnoses to using interactive troubleshooting—and in so doing, many of the engineers who interact with Gandalf have had their opinions on it transformed as well, evolving from skeptics to advocates.

In many Azure services, Gandalf has become a default baseline for all release validations, and it’s exciting to hear how much our on-call engineers trust Gandalf.

Source: azure.microsoft.com

Thursday, 10 June 2021

Advancing in-datacenter critical environment infrastructure availability

Azure Exam Prep, Azure Certification, Azure Exam Prep, Azure Learning, Azure Preparation

There are many factors that can affect critical environment (CE) infrastructure availability—the reliability of the infrastructure building blocks, the controls during the datacenter construction stage, effective health monitoring and event detection schemes, a robust maintenance program, and operational excellence to ensure that every action is taken with careful consideration of related risk implications.

Read More: MS-700: Managing Microsoft Teams

In addition to leading the industry in designing, constructing, commissioning, and operating a high availability CE infrastructure, Microsoft also invests in developing and implementing best practices in the reliability engineering field—aiming to elevate and sustain CE infrastructure availability. In general terms, "reliability engineering" in this blog refers to the engineering discipline that focuses on proactively identifying and assessing time-latent risks which can impact system functionality during its expected useful life—with various failure modes, effects, and mechanisms. As one simple example, when a datacenter is cooled using more energy-efficient free air cooling, we make sure that it’s the CE infrastructure that will deliver a controlled temperature, humidity, and air quality environment. One of the purposes is to prevent corrosion-related failures since electrical shorts or opens can occur under inappropriate conditions such as for the printed circuit board assembly.

The CE for hyperscale cloud datacenters is typically constructed with a blend of off-the-shelf components from a diverse vendor base—equipment like uninterruptible power supplies (UPS), power distribution units (PDU), automatic transfer switches (ATS), air handling units (AHU), and generators. This multi-vendor, multi-technology environment introduces intrinsic complexity as well as variations in robustness, mostly dependent on vendor competencies and experience. The monitoring of the CE infrastructure health is also limited, for example, due to insufficient information about these off-the-shelf components. Monitoring therefore mostly relies on high-level “failure effect” detection, such as from an electrical power management system (EPMS) or a building automation system (BAS).

However, failure effect-based detection schemes generally make it challenging to detect and pinpoint the offending element quickly, thus potentially delaying the rapid recovery of services when failures do occur. The scale of global operations adds further challenges—as variations in local conditions can result in different stresses on the equipment—posing challenges for the effectiveness of the typically time-based preventive maintenance approach. Since some geographic areas experience more notable changes in temperature, humidity, or air quality, a time-based maintenance approach could become wasteful or miss the opportunity to “refresh” system reliability if it does not account for the rate of degradation under the corresponding environmental stress conditions.

Investing in CE reliability engineering

To address these kinds of challenges, Microsoft has invested in establishing the CE reliability engineering function. Many of the related reliability engineering and technologies have emerged from the electronics industry in the last few decades, and there are ample opportunities to adapt, expand, and innovate reliability best practices and solutions for the datacenter CE infrastructure. For example, for off-the-shelf CE components, we have sought a close partnership with CE vendors in application-specific reliability risk evaluations during our equipment selection and qualification stages. During the operation stage, we also closely collaborate with our vendor base to drive continuous improvement through in-depth physical and data analytics—ranging from root cause analyses for individual cases to fleet-wide health behavior data analytics. These partnerships help both Microsoft and our vendors to have clear, data-driven visibility into the related reliability trends, areas of potential risks, and underlying contributing factors, all so that effective resolutions can be introduced, often proactively before more consequential impact occurs. While improving the understanding of potential CE infrastructure risks, reliability engineering is also focused on Research and Development (R&D) efforts that can result in more proactive and effective infrastructure health sensing, such as by capturing parameters correlated with the failure modes and mechanisms for high criticality areas. Internal, as well as external partnerships, have been established in the R&D of relevant approaches, from data mining and machine learning to Physics-of-Failure (PoF) methodologies.

Azure Exam Prep, Azure Certification, Azure Exam Prep, Azure Learning, Azure Preparation
As Microsoft continues to introduce innovations into the CE space, reliability engineering is playing a central role in ensuring the robustness of these solutions by driving both designed-in and built-in reliability through early-stage risk analysis and mitigations. For instance, the Microsoft Sphere-based IoT solution is developed to acquire data securely from the mechanical and electrical power CE system. Reliability engineering closely works with internal and external product design and manufacturing partners in applying both analytical- and PoF-based testing approaches throughout the solution concept, prototype, design, process development, and product deployment stages. One case in point is the concern about electronics’ packaging-related defects, during its manufacture or assembly, or its usage lifecycle. A simulation tool based on finite element analysis (FEA) was employed to identify thermal-mechanical stress points, even when the design is still just a drawing on paper, as these stress points can result in failures within the expected useful life. These points are then closely tracked and characterized (for example, with strain gauges) during environmental stress tests or during manufacturing process steps that can introduce the related thermo-mechanical stresses. Even if the system may still be functional after experiencing these stresses, samples are physically cross-sectioned in critical junctions to identify any early initiation of defects. These in-depth analyses enable concurrent product or process design changes to eliminate the failure likelihood, thereby elevating eventual system availability.

Similar design-for-excellence (DFX) strategies are also being explored for the complex CE infrastructure itself to enable proactive risk identification and prevention opportunities—and prior to the physical deployment of the infrastructure wherever feasible. These CE-related investments and technology advances in the reliability engineering discipline will help Microsoft’s CE infrastructure to build additional robustness mechanisms, in order to meet our customer’s availability expectations for a world-class cloud computing service.

Source: microsoft.com