Showing posts with label Virtual Machines. Show all posts
Showing posts with label Virtual Machines. Show all posts

Saturday, 24 June 2023

Deploy a holistic view of your workload with Azure Native Dynatrace Service

Microsoft and Dynatrace announced the general availability of Azure Native Dynatrace Service in August 2022. The native integration enables organizations to leverage Dynatrace as a part of their overall Microsoft Azure solution. Users can onboard easily to start monitoring their workloads by deploying and managing a Dynatrace resource on Azure.

Azure Native integration enables you to create a Dynatrace environment like you would create any other Azure resource. One of the key advantages of this integration is the ability to seamlessly ship logs and metrics to Dynatrace. By leveraging Dynatrace OneAgent, users can also gather deeper observability data from compute resources such as virtual machines and Azure App Services. This comprehensive data collection ensures that organizations have a holistic view of their Azure workloads and can proactively identify and resolve issues. 

Furthermore, the integration unifies billing for Azure services, including Dynatrace. Users receive a single Azure bill that encompasses all the services consumed on the platform, providing a unified and convenient billing experience. 

Since its release, Dynatrace Service has seen continuous enhancements. In the following sections, we will explore some of the newer capabilities that have been added to further empower organizations in their monitoring and observability efforts. 

Automatic shipping of Azure Monitor platform metrics 


One of the significant advancements during the general availability of Azure Native Dynatrace Service was the automatic forwarding of logs from Azure Monitor to Dynatrace. The log forwarding capability allows you to configure and send Azure Monitor logs to Dynatrace. Logs start to flow to your Dynatrace environment as soon as the Dynatrace resource on Azure is deployed. The Azure experience allows you to view the summary of all the resources being monitored in your subscription. 

Building further, we have now added another key improvement and that is the ability to automatically obtain metrics from the Azure Monitor platform. This enhancement enables users to effectively view the metrics of various services within Azure on the Dynatrace portal. 

To enable metrics collection, customers can simply check a single checkbox on the Azure portal. This streamlined process makes it easy for organizations to start gathering valuable insights. For further customization, users have the option to specify tags to include or exclude specific resources for metric collection. This allows for a more targeted monitoring approach based on specific criteria.

Azure Native Dynatrace Service, Azure Certification, Azure Guides, Azure Prep, Azure Preparation, Azure Tutorial and Materials

The setup of credentials required for the interaction between Dynatrace and Azure is automated, eliminating the need for manual configuration. Once the metrics are collected, users can conveniently view and analyze them on the Dynatrace portal, providing a comprehensive and centralized platform for monitoring and observability. 

Azure Native Dynatrace Service, Azure Certification, Azure Guides, Azure Prep, Azure Preparation, Azure Tutorial and Materials

Together with logs and metrics monitoring capabilities, Azure Native Dynatrace Service provides holistic monitoring of your Azure workloads. 

Native integration availability in new Azure regions 


During general availability, Azure Native Dynatrace Service was available in two regions, the Eastern United States and Western Europe. However, to cater to the growing demand, native integration is now available in additional regions. You can now create a Dynatrace resource in—The United Arab Emirates North (Middle East), Canada Central, and the Western United States—bringing the total number of supported regions to five. You can select the region in the resource creation experience. When selecting a region to provision a Dynatrace resource, the corresponding Dynatrace environment is provisioned in the same Azure region. This ensures that your data remains within the specified region. Hence, it gives you the power to leverage the power of Dynatrace within the Azure region while complying with the specific data residency regulations and preferences of your organization. 

Monitor activity with Azure Active Directory logs


In the realm of cloud business, early detection of security threats is crucial to safeguarding business operations. Azure Active Directory (Azure AD) activity logs—encompassing audit, sign-in, and provisioning logs—offer organizations essential visibility into the activities taking place within their Azure AD tenant. By monitoring these logs, organizations can gain insights into user and application activities, including user sign-in patterns, application changes, and risk activity detection. This level of visibility empowers organizations to respond swiftly and effectively to potential threats, enabling proactive security measures and minimizing the impact of security incidents on their operations. 

With Azure Native Dynatrace Service, you can route your Azure AD logs to Dynatrace by setting Dynatrace as a destination in Azure AD diagnostic settings.

Azure Native Dynatrace Service, Azure Certification, Azure Guides, Azure Prep, Azure Preparation, Azure Tutorial and Materials

Committed to collaboration and integration


The Azure Native integration for Dynatrace has simplified the process of gaining deep insights into workloads. This integration empowers organizations to optimize their resources, enhance application performance, and deliver high availability to their users. Microsoft and Dynatrace remain committed to collaborating and improving the integration to provide a seamless experience for their joint customers. By working together, both companies strive to continually enhance the monitoring and observability capabilities within the Azure ecosystem. 

The product is constantly evolving to deepen the integration, aiming to monitor a wide range of Azure workloads and uplift user convenience throughout the experience. 

Source: microsoft.com

Tuesday, 6 June 2023

Increase gaming performance with NGads V620-series virtual machines

Virtual Machines, Microsoft Career, Microsoft Skills, Microsoft Jobs, Microsoft Prep, Microsoft Preparation, Microsoft Tutorial and Materials, Microsoft Preparation Exam

Gaming customers across the world tend to look for the same critical components when choosing their playing environment: Performance, Affordability, and Timely Content. And for gaming in the cloud, there’s a fourth: Reliability.

With these clear guidelines in mind, we are excited to announce the public preview of our new NGads V620-series virtual machines (VMs). This VM series has GPU, CPU, and memory resources balanced to generate and stream high-quality graphics for a performant, interactive gaming experience hosted on Microsoft Azure. The new NGads instances give online gaming providers the power and stability that they need, at an affordable price.   

The NGads V620-series are GPU-enabled virtual machines powered by AMD Radeon PRO V620 GPU and AMD EPYC 7763 CPUs. The AMD Radeon PRO V620 GPUs have a maximum frame buffer of 32GB which can be divided up to 4 ways through hardware partitioning, or by providing multiple users with access to shared, session-based operating systems such as Windows Server 2022 or Windows 11 EMS. The AMD EPYC CPUs have a base clock speed of 2.45 GHz and a boost speed of 3.5 GHz. VMs are assigned full cores instead of threads, enabling full access to AMD’s powerful Zen 3 cores.

NGads instances come in four sizes, allowing customers to right-size their gaming environments for the performance and cost that best fits their business needs.

The two smallest instances rely on industry-standard SR-IOV technology to partition the GPUs into one-fourth and one-half instances, enabling customers to run workloads with no interference or security concerns between users sharing the same physical graphics card.

The VMs also feature the AMD Software Cloud Edition, which targets the same optimizations available in the consumer gaming version of the Adrenaline driver but is further tested and optimized for the cloud environment.

Instance Configs vCPU (Physical Cores)  GPU Memory (GiB) GPU Partition Size  Memory (GiB)  Azure Network (Gbps) 
Standard_NG8ads_V620_v1 8 ¼ GPU 16 10
Standard_NG16ads_V620_v1  16  16  ½ GPU  32  20 
Standard_NG32ads_V620_v1  32  32  1x GPU  64  40 
Standard_NG32adms_V620_v1  32  32  1x GPU  176   40 

The NGads V620-series VMs will support a new AMD Cloud Software driver that comes in two editions: A Gaming driver with regular updates to support the latest titles, as well as a Professional driver for accelerated Virtual Desktop environments, with Radeon PRO optimizations to support high-end workstation applications.

Microsoft Azure, do more with less


Deployment in Azure enables gaming and desktop providers to take advantage of the infrastructure investments put in place by Microsoft in data centers across the world. This gives our customers the ability to only pay for what they use. They can depend on an infrastructure framework that is constantly kept up to date with highly reliable uptime. Customers can innovate faster to differentiate their offerings and provide customers with a richer experience. As our customers’ business needs expand, they can benefit from the economies of scale available from Azure. In addition, customers can build a more complete and robust solution through integration with the broad range of cloud services for storage, networking, and application management available as part of the Azure offerings.

Flexible workloads, flexible costs


High-performance GPU-accelerated workloads have always ranged from workstation design apps to VDI and simulation rendering. Each of these has the potential to tax even powerful graphics boards. Gaming workloads bring the additional challenges of requiring very fast graphics remoting—the interactive transfer of graphics and user controls over the internet. Further, there is a wide variety of games, connection types, and resolutions available to the user.

The NGads V620-series helps resolve these challenges by providing support for a range of visualization applications so that gaming or desktop service providers can optimize for precisely the experiences expected by the end users. Service provider customers can choose the right-sized VM that will best serve their needs without over-allocating resources. As the needs of their offering change, the common software support across VMs allows service providers to shift to a VM size with either a higher or lower GPU partition, or to shift capacity to other regions of the world as their business footprint expands.

Performance powered by AMD GPU and CPU


The NGads V620-series combines AMD Radeon™ GPU and Epyc™ CPU technology to provide a powerful and well-balanced environment for hosting rich and highly-interactive cloud services. 

The AMD Radeon PRO V620 GPU is based on AMD’s RDNA™ 2 Architecture, AMD Software, and AMD Graphics Virtualization technology. 

Each AMD Radeon PRO V620 GPU is equipped with 32MB of GDDR6 dedicated memory, a 256-bit memory interface with up to 512GB/s bandwidth, and ECC support for data correction.  To enhance the user experience, they are designed with hardware raytracing using 72 Ray Accelerators, 4608 Stream Processors, and a peak Engine Clock of 2200 MHz.

The AMD software supports the DirectX® 12.0, OpenGL®4.6, OpenCL™ 2.2, and Vulkan® 1.1 APIs for broad compatibility with gaming and graphics applications.  This enables the NG series VMs to support a very broad range of workloads from cloud gaming, GPU-enhanced VDI, and GPU-intensive Workstation-as-a-Service solutions.

The NGads V620-series uses GPU Partitioning to virtualize the GPU and provide partitions from the full 32 GB memory size (1x GPU), 16GB (one-half GPU), or 8GB (one-fourth GPU).  The Azure GPU Partitioning is based on the PCIe standard SR-IOV extension, which provides a highly predictable and secure method to host multiple independent user environments on the same hardware GPU board.

The AMD EPYC 7763 CPU is built on the 7nm process technology, featuring AMD Zen 3 cores, Infinity Architecture, and the AMD Infinity Guard suite of security features. The AMD EPYC CPUs have a base clock speed of 2.45GHz and a boost clock speed of 3.5 GHz to allow the user to take advantage of a single powerful core when required by the application.

Source: microsoft.com

Saturday, 6 May 2023

Preparing for future health emergencies with Azure HPC

Azure Exam, Azure Exam Prep, Azure Tutorial and Material, Azure Learning, Azure Certification, Azure Guides, Azure

A once-in-a-century global health emergency accelerates worldwide healthcare innovation and novel medical breakthroughs, all supported by powerful high-performance computing (HPC) capabilities.

COVID-19 has forever changed how nations function in the globally interconnected economy. To this day, it continues to affect and shape how countries respond to health emergencies. COVID-19 has demonstrated just how interconnected our society is and how risks, threats, and contagions can have global implications for many aspects of our daily lives.

COVID-19 was the largest global health emergency in over a century, with nearly 762 million cases reported as of the end of March 2023, according to the World Health Organization. The National Centre for Biotechnology Information points out the frequency and breath of new variants that continues to emerge at regular intervals. In response to this intricate health crisis, the global healthcare community quickly mobilized to better understand the virus, learn its behavior, and work toward preventative treatment measures to minimize the damage to lives across the world. Globally, nations mobilized resources for frontline workers, offered social protection to those most severely affected, and provided vaccine access for the billions who need it.

Recent technological innovations have provided the medical community with access to capabilities, such as HPC, that equipped healthcare professionals to better study, understand, and respond to COVID-19. Globally, healthcare innovators could access unprecedented computing power to design, test, and develop new treatments, faster, better, and more iteratively, than ever before.

Today, Azure HPC enables researchers to unleash the next generation of healthcare breakthroughs. For example, the computational capabilities offered by the Azure HPC HB-series virtual machines, powered by AMD EPYCTM CPU cores, allowed researchers to accelerate insights and advances into genomics, precision medicine, and clinical trials, with near-infinite high-performance bioinformatics infrastructure capabilities.

Since the beginning of COVID-19, companies have been leveraging Azure HPC to develop new treatments, run simulations, and testing at scale—all in preparation for the next health emergency. Azure HPC is helping companies unleash new treatments and health cure capabilities that are ushering in the next generation of treatments and healthcare capabilities, across the entire industry.

High-performance computing making a difference


A leading immunotherapy company partnered with Microsoft to leverage the capabilities of Azure HPC’s high-performance computing, in order to perform detailed computational analyses of the spike protein structure of SARS-CoV-2. Due to the critical nature of the spike protein structure and the role it plays in allowing the invasion of human cells, targeting it for study, analyses, and insights, is a crucial step in the development of treatments to combat the virus.

The company’s engineers and scientists collaborated with Microsoft, and quickly deployed HPC clusters on Azure, containing over 1250 core graphic processing units (GPUs). These GPUs are specifically designed for machine learning and similarly intense computational applications. The Azure HPC clusters augmented the company’s existing GPU clusters—which was already optimized for molecular modelling of proteins, antibodies, and antivirals—bringing a truly high-powered scaled engagement approach to fruition.

By collaborating with Microsoft in this way and making use of the massive, networked computing capabilities and advanced algorithms enabled by Azure HPC, the company was able to generate working models in days rather than the months it would have taken by following traditional approaches.

The incredible amount of computing power will help bolster drug discovery and therapeutic developments. By joining forces and bringing together the incredible power of Azure HPC and cutting edge immunotherapies, it helped contribute to the development of models that allowed researchers to better understand the virus, find novel binding sites to fight the virus, and ultimately guide the development of future treatments and vaccines for the virus.

Powering pharmaceutical research and innovation


The healthcare industry is making remarkable strides in the development of cutting-edge treatments and innovations that are geared towards solving some of the world's greatest healthcare challenges.

For example, researchers are leveraging HPC to transform their research and development effort as well as accelerating the development of new life-saving treatments.

Azure Exam, Azure Exam Prep, Azure Tutorial and Material, Azure Learning, Azure Certification, Azure Guides, Azure

Using a technique producing amorphous solid dispersions (ASD), drug researchers break up active pharmaceutical ingredients and blend them with organic polymers to improve the dissolution rate, bioavailability, and solubility of drug delivery systems. Although a wonder of modern medicine, it is a highly complicated, often lab-based process that can take months.

Swiss-based Molecular Modelling Laboratory (MML), a leader in ASD screening, wanted to pivot its drug research and development to small organic and biomolecular polymers. This approach determines ASD stability prior to formulation, reveals new ASD combinations, enhances drug safety, and helps reduce drug development costs as well as delivery times.

MML chose to leverage Azure HPC resources on more than 18,000 Azure HBv2 virtual machines and to optimize high-throughput drug screening and active pharmaceutical ingredient solubility limit detection, with the aim to alleviate common development hurdles.

The adoption of Azure HPC has helped MML shift from a small start-up to an established business working with some of the top pharmaceutical companies in the world—all in a very short time.

For the global healthcare community, the computational power and scalability of Azure HPC presents an unprecedented opportunity to accelerate pharmaceutical, medical, as well as health innovation. Azure HPC will continue playing a leading role in supporting the healthcare industry to respond optimally to any future global health emergency that may arise.

Source: microsoft.com

Tuesday, 14 March 2023

Azure previews powerful and scalable virtual machine series to accelerate generative AI

Azure, Generative AI, Azure Career, Azure Skills, Azure Jobs, Azure Prep, Azure Preparation, Azure Tutorial and Materials, Azure Guides, Azure

Delivering on the promise of advanced AI for our customers requires supercomputing infrastructure, services, and expertise to address the exponentially increasing size and complexity of the latest models. At Microsoft, we are meeting this challenge by applying a decade of experience in supercomputing and supporting the largest AI training workloads to create AI infrastructure capable of massive performance at scale. The Microsoft Azure cloud, and specifically our graphics processing unit (GPU) accelerated virtual machines (VMs), provide the foundation for many generative AI advancements from both Microsoft and our customers.

“Co-designing supercomputers with Azure has been crucial for scaling our demanding AI training needs, making our research and alignment work on systems like ChatGPT possible.”—Greg Brockman, President and Co-Founder of OpenAI. 

Azure's most powerful and massively scalable AI virtual machine series


Today, Microsoft is introducing the ND H100 v5 VM which enables on-demand in sizes ranging from eight to thousands of NVIDIA H100 GPUs interconnected by NVIDIA Quantum-2 InfiniBand networking. Customers will see significantly faster performance for AI models over our last generation ND A100 v4 VMs with innovative technologies like:

◉ 8x NVIDIA H100 Tensor Core GPUs interconnected via next gen NVSwitch and NVLink 4.0
◉ 400 Gb/s NVIDIA Quantum-2 CX7 InfiniBand per GPU with 3.2Tb/s per VM in a non-blocking fat-tree network
◉ NVSwitch and NVLink 4.0 with 3.6TB/s bisectional bandwidth between 8 local GPUs within each VM
◉ 4th Gen Intel Xeon Scalable processors
◉ PCIE Gen5 host to GPU interconnect with 64GB/s bandwidth per GPU
◉ 16 Channels of 4800MHz DDR5 DIMMs

Delivering exascale AI supercomputers to the cloud


Generative AI applications are rapidly evolving and adding unique value across nearly every industry. From reinventing search with a new AI-powered Microsoft Bing and Edge to AI-powered assistance in Microsoft Dynamics 365, AI is rapidly becoming a pervasive component of software and how we interact with it, and our AI Infrastructure will be there to pave the way. With our experience of delivering multiple-ExaOP supercomputers to Azure customers around the world, customers can trust that they can achieve true supercomputer performance with our infrastructure. For Microsoft and organizations like Inflection, NVIDIA, and OpenAI that have committed to large-scale deployments, this offering will enable a new class of large-scale AI models.

"Our focus on conversational AI requires us to develop and train some of the most complex large language models. Azure's AI infrastructure provides us with the necessary performance to efficiently process these models reliably at a huge scale. We are thrilled about the new VMs on Azure and the increased performance they will bring to our AI development efforts."—Mustafa Suleyman, CEO, Inflection.

AI at scale is built into Azure’s DNA. Our initial investments in large language model research, like Turing, and engineering milestones such as building the first AI supercomputer in the cloud prepared us for the moment when generative artificial intelligence became possible. Azure services like Azure Machine Learning make our AI supercomputer accessible to customers for model training and Azure OpenAI Service enables customers to tap into the power of large-scale generative AI models. Scale has always been our north star to optimize Azure for AI. We’re now bringing supercomputing capabilities to startups and companies of all sizes, without requiring the capital for massive physical hardware or software investments.

“NVIDIA and Microsoft Azure have collaborated through multiple generations of products to bring leading AI innovations to enterprises around the world. The NDv5 H100 virtual machines will help power a new era of generative AI applications and services.”—Ian Buck, Vice President of hyperscale and high-performance computing at NVIDIA. 

Today we are announcing that ND H100 v5 is available for preview and will become a standard offering in the Azure portfolio, allowing anyone to unlock the potential of AI at Scale in the cloud.

Source: azure.microsoft.com

Thursday, 22 December 2022

Microsoft Innovation in RAN Analytics and Control

Currently, Microsoft is working on RAN Analytics and Control technologies for virtualized RAN running on Microsoft Edge platforms. Our goal is to empower any virtualized RAN solution provider and operators to realize the full potential of disaggregated and programmable networks. We aim to develop platform technologies that virtualized RAN vendors can leverage to gain analytics insights in their RAN software operations, and to use these insights for operational automations, machine learning, and AI-driven optimizations.

Microsoft has recently made important progress in RAN analytics and control technology. Microsoft Azure for Operators is introducing flexible, dynamically loaded service models to both the RAN software stack and cloud/edge platforms hosting the RAN, to accelerate the pace of innovation in Open RAN.

The goal of Open RAN is to accelerate innovation in the RAN space through the disaggregation of functions and exposure of internal interfaces for interoperability, controllability, and programmability. The current standardization effort of O-RAN by O-RAN Alliance, specifies the RAN Intelligent Controller (RIC) architecture that exposes a set of telemetry and control interfaces with predefined service models (known as the E2 interface). Open RAN vendors are expected to implement all E2 service models specified in the standard. Near-real-time RAN controls are made possible with xApp applications accessing these service models.

Microsoft’s innovation extends this standard-yet-static interface. It introduces the capability of getting detailed internal states and real-time telemetric data out of the live RAN software in a dynamic fashion for new RAN control applications. With this technology, together with detailed platform telemetry, operators can achieve better network monitoring and performance optimization for their 5G networks, and enable new AI, analytics, and automation capabilities that were not possible before.

This year, Microsoft, together with contributions from Intel and Capgemini, has developed an analytics and control approach that was recognized with the Light Reading Editor’s Choice award under the category of Outstanding Use case: Service provider AI. This innovation calls for dynamic services models for Open RAN.

Dynamic service models for real-time RAN control


There are many RAN control use cases that require dynamic service models beyond those specified in O-RAN today, such as access to IQ samples, RLC and MAC queue sizes, and packet retransmission information. These high-volume real-time data need to be aggregated and compressed before being delivered to the xApp. Also, detailed data from different RAN modules across different layers like L1, L2, and L3 may need to be collected and correlated in real-time before any useful insight can be derived and shared with xApp. Further, a virtualized RAN offers so many more possibilities, that any static interface or service model may be ineffective in meeting the more advanced real-time control needs.

One such example occurs with interference detection. Today, operators typically need to do a drive test to detect external interference in a macro cell. But now, Open RAN has the potential to replace the expensive truck roll with a software program that detects interference signals at the RAN’s L1 layer. However, this will require a new data service model with direct access to raw IQ samples at the physical layer. Another example exists in dynamic power saving. If a RAN power controller can see the number of packets queued at various places in the live RAN system, then it can estimate the pending process loads and optimize the CPU frequency at a very high pace, in order to reduce the RAN server power consumption. Our study has shown that we can reduce the RAN power consumption by 30 percent through this method—even during busy periods. To support this in Open RAN, we will need a new service model that exposes packet queuing information.

These new use cases are envisioned for the time after the current E2 interface has been standardized. To achieve them, though, we need new RAN platform technologies to quickly extend this interface to support these and future advanced RAN control applications.

The Microsoft RAN analytics and control framework


The Microsoft RAN analytics and control framework extends the current RIC service models in O-RAN architecture to be both flexible and dynamic. In the process, the framework allows RAN solution providers and operators to define their own service models for dynamic RAN monitoring and control. Here, the underlying technology is a runtime system that can dynamically load and execute third-party code in a trusted and safe manner.

This system enables operators and trusted third-party developers to write their own telemetry, control, and inference pieces of code (called “codelets”) that can be deployed at runtime at various points in the RAN software stack, without disrupting the RAN operations. The codelets are executed inline in the live RAN system and on its critical paths, allowing them to get direct access to all important internal raw RAN data structures, to collect statistics, and to make real-time inference and control decisions.

To ensure security and safety, the codelets checked with static verified with verification tools before they can be loaded, and they will be automatically pre-empted if running longer than the predefined execution budgets. The dynamic code extension system is the same as the Extended Berkeley Packet Filter (eBPF), which is a proven technology that has been entrusted to run custom codes in Linux kernels on millions of mission-critical servers around the globe. The inline execution is also extremely fast, typically incurring less than one percent of overhead on the existing RAN operations.

The following image illustrates the overall framework and the dynamic service model denoted by the star circle with the letter D.

Microsoft Innovation, RAN Analytics, Microsoft Career, Microsoft Skills, Microsoft Jobs, Microsoft Tutorial and Materials, Microsoft Certification, Microsoft Guides

The benefit of the dynamic extension framework with low-latency control is that it can open the opportunity for third-party real-time control algorithms. Traditionally, due to the tight timing constraint, a real-time control algorithm must be tightly implemented and integrated inside the RAN system. The Microsoft RAN analytics framework allows RAN software to delegate certain real-time control to RIC, potentially leading to a future marketplace where real-time control algorithms, machine learning, and AI models for optimizations may be possible.

Microsoft, Intel, and Capgemini have jointly prototyped this technology in Intel’s FlexRAN™ reference software and Capgemini’s 5G RAN. We have also identified standard instrumentation points aligned with the standard 3GPP RAN architecture to achieve higher visibility into the RAN’s internal state. We have further developed 17 dynamic service models, and enabled many new and exciting applications that were previously not thought possible.

Examples of new applications of RAN analytics


With this new Analytics and Control Framework, applications of dynamic power savings and interference detection described earlier can now be realized.

RAN-agnostic dynamic power saving

5G RAN energy consumption is a major OPEX item for any mobile operator. As a result, it is paramount for a RAN platform provider to find any opportunity to save power when running the RAN software. One such opportunity can be found by stepping down the RAN server CPU frequency when the RAN processing load is not at full capacity. This is indeed promising because internet traffic is intrinsically “bursty”; even during peak hours, the network is rarely operated at full capacity.

However, any dynamic RAN power controller must also have accurate load prediction and fast reaction in millisecond timescale. Otherwise, if one part of RAN is in hibernation, then any instant traffic burst will cause serious performance issues, or even crashes. The Microsoft RAN analytics framework with dynamic service models and low-latency control-loop makes it possible to write a novel CPU frequency prediction algorithm based on the number of active users, and changes in different queue sizes. We have implemented this algorithm on top of Capgemini 5G RAN and Intel FlexRAN™ reference software, and we achieved up to 30 percent energy savings—even during busy periods.

Interference detection

External wireless interference has long been a source of performance issues in cellular networks. Detecting external wireless interference is difficult and often requires a truck roll with specialized equipment and experts to detect it. With dynamic service models, we can turn an O-RAN 5G base station into a software-defined radio that can detect and characterize external wireless interference without affecting the radio performance. We have developed a dynamic service model that averages the received IQ samples across frequency chunks and times inside an L1 of the FlexRAN™ reference software stack. The service model in turn reports the averages to an application that runs an AI and machine learning model for anomaly detection, in order to detect when the noise floor increases.

Virtualized and software-based RAN solution offer immense potential of programmable networks that can leverage AI, machine learning, and analytics to improve network efficiency. Dynamic service models for O-RAN interfaces further enhances the pace of innovation with added flexibility and security.

Source: microsoft.com

Saturday, 3 December 2022

Azure comes to Dallas for Supercomputing

Azure Exam, Azure Exam Prep, Azure Tutorial and Materials, Azure Career, Azure Skill, Azure Job

Supercomputing (SC), held annually, is arguably one of the biggest annual events in high-performance computing (HPC). It’s a great chance for the community to connect and learn from one another, and for Microsoft, the goal remains the same. After a couple years of virtual-only events, we’re thrilled to be joining you in person in Dallas this year—and we’ve got some exciting things in store.

Step into our Microsoft Booth


This year we’ll be located at booth #2433—make sure to stop by and say "hi." We’ll have several goodies for you to take, a caricature artist to create a one-of-a-kind keepsake of your experience at the event, and two hardware booths to let you see in person some of the technology that powers Microsoft Azure HPC + AI virtual machines (VMs). We’ll even be showcasing our newest product launch in the hardware bar. Feeling tired? Come take a break in our lounge or café area for some much-needed relief, networking, or coffee. And of course, an event wouldn’t be the same without our Microsoft specialists and partners sharing in our booth.

Explore our joint booth with NVIDIA


We’ll also have a joint booth with NVIDIA, located at #2409. We’ll have Microsoft experts and partners giving presentations in this booth to share insights on the confluence of AI with HPC using NVIDIA accelerated computing on Azure. Don’t be shy about stopping by to talk to our subject matter experts, network with peers or simply get off your feet.

Want to take a break from the hustle of the conference and stretch your creative muscles? Come over to our chalkboard and create an image of what “Make AI Your Reality” means to you. Here’s what to do—draw your image, take a picture and post it to your favorite channel with #MakeAIYourReality, and then show it to a Microsoft or NVIDIA representative in the booth to get your raffle ticket for a chance to win an NVIDIA Jetson Nano developer kit—a small but powerful computer to start learning about building AI-enabled applications with ready-to-try projects and community support.

Attend the Women in HPC (WHPC) networking reception


We are tremendously honored to be sponsoring the annual Women in High-Performance Computing networking reception. Not only will this drive awareness of diversity and inclusivity topics, but will produce understandings about how we can all work towards improving the under-representation of women in supercomputing.

Visit the Student Cluster Competition


We are so excited to be supporting the SC Student Cluster Competition once again this year! Every year, both undergraduate and high school students design and build small clusters, learn scientific applications, apply optimization techniques for their chosen architectures, and compete in a 48-hour challenge at the SC event to complete real-world workloads, demonstrating their HPC knowledge for conference attendees and judges.

Catch up on what’s new at this year's Supercomputing


Hear from our customers

With a tight timeline, UD Trucks collaborated with Microsoft and executing partner HCLTech to design and deploy a Microsoft Azure–based system for its previously on-premises simulation and design processes, completing the project in only one month. The new system’s capabilities provide better computational results and have opened new avenues of data innovation, in addition to reducing costs by about 30 percent. The project’s success has led the company to develop plans to move all of its systems to Azure.

Rimac Technology is now running Ubercloud on Microsoft Azure to support engineering simulations during the development and testing of electrical vehicles and components. The move has allowed them to manage greater model complexity and take advantage of increased processing speed and scale.


University of bath moved almost all its existing supercomputing resources to the cloud with Microsoft Azure HPC + AI. To power research workloads, it deployed Azure HPC + AI resources on 21 virtual machine instances, ensuring that the right virtual machines can be spun up for any project.


Kensington has decreased its engine runtime from 20 hours to just 25 minutes—for 1.9 billion total calculations with every run. With the stronger insights and predictions that it gains from its optimized models and results, Kensington creates more tailored products and better serves its customers.

Source: microsoft.com

Thursday, 1 December 2022

Voltus and Azure—no power integrity challenge too big to solve

With the advent of AI and hyperscale designs on advanced nodes, it is common to see designs in over 50 billion transistor categories with tens to 100 billion plus nodes in the on-chip power network. This explosion in scale requires solutions that meet the following requirements:

◉ High performance and capacity.
◉ Elasticity.
◉ Manage varying compute resource requirements.
◉ Low cost to manage the exponential increase in compute requirements.

Voltus on Azure


Voltus is a leading IC Power Integrity Signoff Solution from Cadence Design Systems. It is used by top chip design companies to verify the reliability of their power networks on chip (NoC) and enables power integrity and thermal analysis at the system level.

Microsoft Azure provides a cloud-based high-performance computing (HPC) infrastructure with security, reliability, and scalability that is a natural fit for electronic design automation (EDA) workloads, especially power integrity analysis.

Azure can support both a hybrid model as well as an all-in model. In the hybrid model customers mainly use their on-premises infrastructure but can add to their compute and storage capacity on an on-demand basis to satisfy peak demand. The hybrid approach is typically used by customers new to using the cloud. In an all-in model, customers primarily use Azure infrastructure for all their EDA workloads. The all-in model is a great use case for startups and customers who really want to optimize their costs while taking advantage of the scale and flexibility of Azure. Voltus supports both the hybrid as well as the all-in model with Azure.

Managing variable compute costs through the design cycle


Using Azure can help customers optimize their costs as compute requirements will vary through the design cycle with lower requirements early on and peak demand near signoff. This is in contrast to the high fixed cost of on-premises infrastructure.

Running Voltus on Azure


We have used a block and full Chip test case to demonstrate our results.

Azure, Azure Exam, Azure Tutorial and Material, Azure Skills, Azure Jobs, Azure Prep, Azure Preparation, Azure Chip

The Azure team selected Edsv4 virtual machines (VMs) based on second-generation Intel Xeon Platinum 8272CL (Cascade Lake). These VMs are well suited for both compute and memory-intensive workloads.

The Voltus use case setup on Azure is illustrated in Figure 1.

Azure, Azure Exam, Azure Tutorial and Material, Azure Skills, Azure Jobs, Azure Prep, Azure Preparation, Azure Chip
Figure 1

High performance and elasticity


Voltus has a fully distributed and scalable architecture. Every step of the power integrity analysis flow, from design parsing to the solver, is fully distributed and scalable. Data from each part of the automatically partitioned design is assigned to compute nodes on the compute infrastructure for various steps in the analysis. This process is managed by a master machine as illustrated in Figure 2.

Azure, Azure Exam, Azure Tutorial and Material, Azure Skills, Azure Jobs, Azure Prep, Azure Preparation, Azure Chip
Figure 2

The level of distribution is user-controlled, which allows the user to take advantage of compute elasticity and manage performance. As Figure 3 illustrates for both the block and full chip run, we observe near-linear scalability in performance with respect to the number of CPUs.

Azure, Azure Exam, Azure Tutorial and Material, Azure Skills, Azure Jobs, Azure Prep, Azure Preparation, Azure Chip
Figure 3

Higher performance with lower costs


Believe it or not, that is indeed true. The elasticity of Voltus architecture enables the tool to run faster with a higher number of CPUs and since the CPUs are used for a smaller amount of time, the result is that the total cost drops to an optimal point. This can be seen at both the block and full chip levels as illustrated in Figure 3. This is a win-win situation where you can improve your performance and reduce your costs.

Azure, Azure Exam, Azure Tutorial and Material, Azure Skills, Azure Jobs, Azure Prep, Azure Preparation, Azure Chip

Azure, Azure Exam, Azure Tutorial and Material, Azure Skills, Azure Jobs, Azure Prep, Azure Preparation, Azure Chip
Figure 4

The magic of Voltus hierarchical analysis


Designers can further increase their performance and reduce cost by using Voltus XM hierarchical analysis. With Voltus XM, block-level models can be used instead of the full flattened design as illustrated in Figure 5. This method significantly reduces node count while maintaining accuracy. We can even further reduce our runtime and costs with Voltus XM and Azure. We observe a 4.5x reduction in cost and a 2x improvement in performance over the flat run for the full chip test case (Figure 6).

Azure, Azure Exam, Azure Tutorial and Material, Azure Skills, Azure Jobs, Azure Prep, Azure Preparation, Azure Chip
Figure 5

Azure, Azure Exam, Azure Tutorial and Material, Azure Skills, Azure Jobs, Azure Prep, Azure Preparation, Azure Chip
Figure 6

We have demonstrated the benefit of using Voltus on Azure at both the block level and chip level. These benchmarks show that customers can not only just benefit from higher performance using elastic compute, but there is an optimal point for performance and cost. Using Voltus XM hierarchical analysis further improves cost and performance. With Voltus on Azure, semiconductor companies have the ideal solution to verify power integrity for their most complex designs.

Source: microsoft.com

Saturday, 12 November 2022

Announcing more Azure VMware Solution enhancements

Azure VMware Solution, Azure Career, Azure Skills, Azure Jobs, Azure Preparation, Azure

I’m writing to you today from VMware Explore in Barcelona, where my team and I are presenting to and meeting with customers and partners in person! When we launched Azure VMware Solution two years ago amid a pandemic, IT agility became a top priority as organizations scrambled to enable remote work and ensure business resilience via cloud solutions. In today’s economic climate most organizations want to do more with less. They recognize that by running workloads in the cloud, they can respond more rapidly and reduce IT infrastructure costs.

"I can definitely say that Azure—and in particular Azure VMware Solution—is the right solution for us. It allows us to seamlessly move from on-premises to the cloud, thereby freeing up resources and capital investments that can be used where they are needed more.”—Giorgio Veronesi, Sr. Vice President of ICT Infrastructure, Snam.

Given that TCO is top priority for most companies in the current economic climate, migrating your VMware workloads to Azure is a great way to reduce the cost of maintaining an on-premises VMware environment. Because every customer starts their cloud journey at a different place, we help enable customers to migrate to the cloud on their terms and maintain support for the business platforms and investments they have today.  Azure VMware Solution is an easy way to extend and migrate existing VMware Private Clouds to run them natively on Azure. Azure VMware Solution offers symmetry with on-premises environments, which helps to accelerate datacenter migrations, so customers recognize the benefits of the cloud sooner.

"With help from Microsoft and Mobiz, we were able to deliver a fully qualified landing zone in Azure in one-third the time and at one-third the budget compared to previous cloud efforts."—Sam Chenaur: Vice President and Global Head of Infrastructure, Sanofi.

In keeping with the goal of doing more with less, Microsoft’s unique Azure Hybrid Benefit and Extended Security Updates for Windows Server and SQL Server, Azure VMware Solution is one of the fastest and most cost-effective ways to seamlessly migrate and run VMware in the cloud.

Check out what’s new in Azure VMware Solution


I am excited to share some of the recent updates we’ve made to Azure VMware Solution.

◉ Stretched Clusters for Azure VMware Solution, now in preview, provides 99.99 percent uptime for mission critical applications that require the highest availability. In times of availability zone failure, your virtual machines (VMs) and applications automatically failover to an unaffected availability zone with no application impact.

◉ Azure NetApp Files Datastores is now generally available to run your storage intensive workloads on Azure VMware Solution. This integration between Azure VMware Solution and Azure NetApp Files enables you to create datastores via the Azure VMware Solution resource provider with Azure NetApp Files NFS volumes and attach the datastores to your private cloud clusters of choice.

◉ Customer-managed keys for Azure VMware Solution is now in preview, both supporting higher security for customers’ mission-critical workloads and providing you with control over your encrypted vSAN data on Azure VMware Solution. With this feature, you can use Azure Key Vault to generate customer-managed keys as well as centralize and streamline the key management process.

◉ New node sizing for Azure VMware Solution. Start leveraging Azure VMware Solution across two new node sizes with the general availability of AV36P and AV52 in AVS. With these new node sizes organizations can optimize their workloads for memory and storage with AV36P and AV52.

◉ Microsoft Azure native services let you monitor, manage, and protect your virtual machines (VMs) in a hybrid environment (Azure, Azure VMware Solution, and on-premises). Here are some of the existing Azure services: Azure Arc, Azure Monitor, Microsoft Defender for Cloud, Azure Update Management, and Log Analytics Workspace.

Source: microsoft.com

Saturday, 15 October 2022

Ensure zone resilient outbound connectivity with NAT gateway

Our customers—across all industries—have a critical need for highly available and resilient cloud frameworks to ensure business continuity and adaptability of ever-growing workloads. One way that customers can achieve resilient and reliable infrastructures in Microsoft Azure (for outbound connectivity) is by setting up their deployments across availability zones in a region.

When customers need to connect outbound to the internet from their Azure infrastructures, Network Address Translation (NAT) gateway is the best way. NAT gateway is a zonal resource that is configured to subnets from the same virtual network, which means that it can be deployed to individual zones to allow outbound connectivity. Subnets and virtual networks, on the other hand, are regional constructs that are not restricted to individual zones. Subnets can contain virtual machine instances or scale sets spanning across multiple availability zones.

Even without being able to traverse multiple availability zones, NAT gateway still provides a highly resilient and reliable way to connect outbound to the internet. This is because it does not rely on any single compute instance like a virtual machine. Instead, NAT gateway leverages software-defined networking to operate as a fully managed and distributed service with built-in redundancy. This built-in redundancy means that customers are unlikely to experience individual NAT gateway resource outages or downtime in their Azure infrastructures.

To ensure that you have the optimal outbound configuration to meet your availability and security needs while also safeguarding against zonal outages, let’s look at how to create zone resilient setups in Azure with NAT gateway.

Zone resilient outbound connectivity scenarios with NAT gateway


Customer setup

Let's say you are a retailer who is preparing for an upcoming Black Friday event. You anticipate that traffic to your retail website will increase significantly on the day of the sale. You decide to deploy a virtual machine scale set (VMSS) so that way your compute resources can automatically scale out to meet the increased traffic demands. Scalability is not the only requirement you have in preparation for this event, but also resiliency and security. To ensure that you safeguard against potential zonal outages that could impact traffic flow, you decide to deploy these VMSS across multiple availability zones. In addition to using VMSS in multiple availability zones, you plan to use NAT gateway to handle all outbound traffic flow in a scalable, secure, and reliable manner.

How should you set up your NAT gateway with your VMSS across multiple availability zones? Let’s take a look at a few different configurations along with which setups will and won’t work.

Scenario 1: Set up a single zonal NAT gateway with your zone-spanning VMSS

First, you decide to deploy a single NAT gateway resource to availability zone 1 and your VMSS across all three availability zones within the same subnet. You then configure your NAT gateway to this single subnet and to a /28 public IP prefix, which provides you a contiguous set of 16 public IP addresses for connecting outbound. Does this setup safeguard you against potential zone outages? No.

Figure 1: A single zonal NAT gateway configured to a zone-spanning set of virtual machines does not provide optimal zone resiliency. NAT gateway is deployed out of zone 1 and configured to a subnet that contains a VMSS that spans across all three availability zones of the Azure region. If availability zone 1 goes down, outbound connectivity across all three zones will also go down.

Here’s why:

1. If the zone that goes down is also the zone in which NAT gateway has been deployed then all outgoing traffic from virtual machines across all zones will be blocked.

2. If the zone that goes down is different than the zone that NAT gateway has been deployed in, then outgoing traffic from the other zones will still occur and only virtual machines from the zone that has gone down will be impacted.

Scenario 2: Attach multiple NAT gateways to a single subnet

Since the previous configuration will not provide the highest degree of resiliency, you decide you will instead deploy 3 NAT gateway resources, one in each availability zone, and attach them to the subnet that contains the VMSS. Will this setup work? Unfortunately, no.

Figure 2: Multiple NAT gateways cannot be attached to a single subnet by design.

Here’s why:

A subnet cannot have more than one NAT gateway attached to it and it is not possible to set up multiple NAT gateways on a single subnet. When NAT gateway is configured to a subnet, NAT gateway becomes the default next hop type for network traffic before reaching the internet. Consequently, virtual machines in a subnet will source NAT to the public IP address(es) of NAT gateway before egressing to the internet. If more than one NAT gateway were to be attached to the same subnet, the subnet would not know which NAT gateway to use to send outbound traffic.

Scenario 3: Deploy zonal NAT gateways with zonally configured VMSS for optimal zone resiliency

What is the optimal solution then for creating a secure, resilient, and scalable outbound setup? The solution is to deploy a VMSS in each availability zone, configure each to their own respective subnet and then attach each subnet to a zonal NAT gateway resource.

Figure 3: Zonal NAT gateways configured to individual subnets for zonal VMSS provide optimal zone resiliency for outbound connectivity.

Deploying zonal NAT gateways to match the zones of the VMSS provides the greatest protection against zonal outages. Should one of the availability zones go down, the other two zones will still be able to egress outbound traffic from the other two zonal NAT gateway resources.

Summary of zone resilient scenarios with NAT gateway


Scenario Description Rating
Scenario 1 Set up a single zonal NAT gateway with your VMSS that spans across multiple availability zones but confined to a single subnet. Not recommended: if the zone that NAT gateway is located in goes down then outbound connectivity for all VMs in the scale set goes down.
Scenario 2  Attach multiple zonal NAT gateways to a subnet that contains zone-spanning virtual machines.  Not possible: multiple NAT gateways cannot be associated to a single subnet by design. 
Scenario 3  Deploy zonal NAT gateways to separate subnets with zonally configured VMSS.  Optimal configuration to provide zone resiliency and protect against outages. 

FAQ on NAT gateway and availability zones


1. What does it mean to have a "no zone" NAT gateway?

◉ "No zone" is the default availability zone selected when you deploy a NAT gateway resource. No zone means that Azure places the NAT gateway resource into a zone for you, but you do not have visibility into which zone it is specifically placed. It is recommended that you deploy your NAT gateway to specific zones so that you know in which zone your NAT gateway resource resides. Once NAT gateway is deployed, the availability zone designation cannot be changed.

2. If I have Load Balancer or instance-level public IPs (IL PIPs) on virtual machines and NAT gateway deployed in the same virtual network and NAT gateway or an availability zone goes down, will Azure fall back to using Load Balancer or IL PIPs for all outbound traffic?

◉ Azure will not failover to using Load Balancer or IL PIPs for handling outbound traffic when NAT gateway is configured to a subnet. After NAT gateway has been attached to a subnet, the user-defined route (UDR) at the source virtual machine will always direct virtual machine–initiated packets to the NAT gateway even if the NAT gateway goes down.

Source: microsoft.com

Tuesday, 4 October 2022

Advancing anomaly detection with AIOps—introducing AiDice

In Microsoft Azure, we invest tremendous efforts in ensuring our services are reliable by predicting and mitigating failures as quickly as we can. In large-scale cloud systems, however, we may still experience unexpected issues simply due to the massive scale of the system. Given this, using AIOps to continuously monitor health metrics is fundamental to running a cloud system successfully, as we have shared in our earlier posts. First, we shared more about this in Advancing Azure service quality with artificial intelligence: AIOps. We also shared an example deep dive of how we use AIOps to help Azure in the safe deployment space in Advancing safe deployment with AIOps. Today, we share another example, this time about how AI is used in the field of anomaly detection. Specifically, we introduce AiDice, a novel anomaly detection algorithm developed jointly by Microsoft Research and Microsoft Azure that identifies anomalies in large-scale, multi-dimensional time series data. AiDice not only captures incidents quickly, it also provides engineers with important context that helps them diagnose issues more effectively, providing the best experience possible for end customers.

Why are AIOps needed for anomaly detection?


We need AIOps for anomaly detection because the data volume is simply too large to analyze without AI. In large-scale cloud environments, we monitor an innumerable number of cloud components, and each component logs countless rows of data. In addition, each row of data for any given cloud component might contain dozens of columns such as the timestamp, the hardware type of the virtual machine, the generation number, the OS version, the datacenter where the nodes hosting the virtual machine stay in, or the country. The structure of the data we have is essentially multi-dimensional time series data, which contains an exponential number of individual time series due to the various combinations of dimensions. This means that iterating through and monitoring every single time series is simply not practical—applying AIOps is necessary.

How did we approach this, before AiDice?


Before AiDice, the way we handled anomaly detection in large-scale, high-dimensional time series data was to conduct anomaly detection on a selected set of dimensions that were the most important. By focusing on a scoped subset, we would be able to detect anomalies within those combinations quickly. Once these anomalies were detected, engineers would then dive deeper into the issues, using pivot tables to drill down into the other dimensions not included to better diagnose the issue. Although this approach worked, we saw two key opportunities to improve the process. First, the old approach required a lot of manual effort by engineers to determine the exact pivot of anomalies. Second, the approach also limited the scope of direct monitoring by only allowing us to input a limited number of dimensions into our anomaly detection algorithms. Given these reasons, Microsoft Research and Azure worked together to develop AiDice, which improves both of these areas.

How do we approach this now with AiDice, and how does it work?


Now with AiDice, we can automatically localize pivots on time series data even if looking at dozens of dimensions at the same time. This allows us to add a lot more attributes, whether that be the hardware generation or hardware microcode, the OS version, or the networking agent version. Though this makes the search space much larger, AiDice encodes the problem as a combinatorial optimization problem, allowing it to search through the space more efficiently than traditional approaches. Brief details of AiDice are described below, but to see a full explanation of the algorithm, please see the paper published at the ESEC/FSE '20: 28th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE 2020).

Part 1: AiDice algorithm—formulation as a search problem


The AiDice algorithm works by first turning the data into a search problem. Search nodes are formed by starting at a given pivot and building the relationships out to the neighbors. For example, if we take a node, "Country=USA, Datacenter=DC1, DiskType=SSD", we can form out the neighboring nodes by swapping, adding, or removing a dimension-value pair, as shown in the diagram below.

AIOps, Microsoft Azure, Microsoft Certification, Microsoft Career, Microsoft Skills, Microsoft Jobs, Microsoft Prep, Microsoft Preparation, Microsoft Tutorial and Materials

Part 2: AiDice algorithm—objective function


Next, the AiDice algorithm searches through the search space in a smart manner by maximizing an objective function that emphasizes two key components. First, the bigger the sudden burst or change in errors, the higher AiDice scores the objective function. Second, the higher the proportion of the errors that occur in this pivot in relation to the total number of errors, the higher AiDice scores the objective function. For example, if there are 5,000 total errors that occurred, it is more important to alert the user about the pivot that went from 3000 errors to 4000 errors than the pivot that went from 10 to 20 errors.

Part 3: Customization of alerts to reduce noise


Next, the alerts that AiDice produces need to be filtered and customized to be less noisy and more actionable since the results so far are optimized from a mathematical perspective but have not yet incorporated domain knowledge around the meaning of the input data. This step can vary widely depending on the nature of the input data, but an example could be that consecutive alerts that share the same error code may be grouped together to reduce the number of total alerts.

AiDice in action—an example


The following is a real example in which AiDice helped detect a real issue early on. The details are altered for confidentiality reasons.

◉ We applied AiDice to monitor low memory error events in a certain type of virtual machine with more than a dozen dimensions of attribute information alongside the fault count, including the region, the datacenter location, the cluster, the build, the RAM, or the event type.

◉ AiDice identified an increase in the number of low memory events on distinct nodes in a particular pivot, which indicated a memory leak.
    ◉ Build=11.11111, Ram=00.0, ProviderName=Xxxxx-x-Xxxxxx, EventType=8888 (details have been altered for privacy).

◉ When looking at the aggregate trend, this issue is hidden and without AiDice it would take manual effort to detect the exact location of the issue (see graphs below, data normalized for privacy).

◉ The engineer responsible for the ticket looked at the alert and some example cases shown in the alerts to quickly able figure out what was going on.

AIOps, Microsoft Azure, Microsoft Certification, Microsoft Career, Microsoft Skills, Microsoft Jobs, Microsoft Prep, Microsoft Preparation, Microsoft Tutorial and Materials

AIOps, Microsoft Azure, Microsoft Certification, Microsoft Career, Microsoft Skills, Microsoft Jobs, Microsoft Prep, Microsoft Preparation, Microsoft Tutorial and Materials

In this real-world example, AiDice was able to detect an issue in a dimension combination that was causing a particular error type in an automatic fashion, quickly and efficiently. Soon after, the memory leak was discovered and Azure engineers were able to mitigate the issue.

Looking forward


Looking ahead, we hope to improve AiDice to make Azure even more resilient and reliable. Specifically, we plan to:

◉ Support additional scenarios in Azure: AiDice is being applied to many scenarios in Azure already, but the algorithm has room to improve with respect to the types of metrics it can operate on. Microsoft Azure and the Microsoft Research team are working together to support more metric scenarios.

◉ Prepare additional data feeds in Azure for AiDice: In addition to upgrading the AiDice algorithm itself to support more scenarios, we are also working to add supporting attributes to certain data sources to fully leverage the power of AiDice.

Source: microsoft.com