Saturday, 12 January 2019

Best practices for alerting on metrics with Azure Database for MariaDB monitoring

Microsoft’s Azure Database for open sources announced the general availability of MariaDB. This blog intends to share some guidance and best practices for alerting on the most commonly monitored metrics for MariaDB.

Whether you are a developer, a database analyst, a site reliability engineer, or a DevOps professional at your company, monitoring databases is an important part of maintaining the reliability, availability, and performance of your MariaDB server. There are various metrics available for you in Azure Database for MariaDB to get insights on the behavior of the server. You can also set alerts on these metrics using the Azure portal or Azure CLI.

Azure Database, Azure MariaDB, Azure Tutorial and Material, Azure Guides, Azure Certification

With modern applications evolving from a traditional on-premises approach to becoming more hybrid or cloud native, there is also a need to adopt some best practices for a successful monitoring strategy on a hybrid/public cloud. Here are some example best practices on how you can use monitoring data on your MariaDB server and areas you can consider improving based on these various metrics.

Active connections


Sample threshold (percentage or value): 80 percent of total connection limit for greater than or equal to 30 minutes, checked every five minutes.

Things to check

If you notice that active connections are at 80 percent of the total limit for the past half hour, verify if this is expected based on the workload.
If you think the load is expected, active connections limits can be increased by upgrading the pricing tier or vCores.

Azure Database, Azure MariaDB, Azure Tutorial and Material, Azure Guides, Azure Certification

Failed connections


Sample threshold (percentage or value): 10 failed connections in the last 30 minutes, checked every five minutes.

Things to check

If you see connection request failures over the last half hour, verify if this is expected by checking the logs for failure reasons.

Azure Database, Azure MariaDB, Azure Tutorial and Material, Azure Guides, Azure Certification

◈ If this is a user error, take the appropriate action. For example, if authentication yields a failed error check your username/password.
◈ If the error is SSL related, check the SSL settings and input parameters are properly configured.
     ◈ Example: psql "sslmode=verify-ca sslrootcert=root.crt host=mydemoserver.mariadb.database.azure.com dbname=mariadb user=mylogin@mydemoserver"

CPU percent or memory percent


Sample threshold (percent or value): 100 percent for five minutes or 95 percent for more than two hours.

Things to check

◈ If you have hit 100 percent CPU or memory usage, check your application telemetry or logs to understand the impact of the errors.
◈ Review the number of active connections. If your application has exceeded the max connections or is reaching the limits, then consider scaling up compute.

IO percent


Sample threshold (percent or value): 90 percent usage for greater than or equal to 60 minutes.

Things to check

◈ If you see that IOPS is at 90 percent for one hour or more, verify if this is expected based on the application workload.
◈ If you expect a high load, then increase the IOPS limit by increasing storage. 

Storage


The storage you provision is the amount of storage capacity available to your Azure Database for PostgreSQL server. The storage is used for the database files, temporary files, transaction logs, and the PostgreSQL server logs. The total amount of storage you provision also defines the I/O capacity available to your server.

Basic General purpose  Memory optimized 
Storage type Azure Standard Storage Azure Premium Storage Azure Premium Storage 
Storage size  5GB TO 1TB  5GB to 4TB  5GB to 4TB 
Storage increment size  1GB  1GB  1GB
IOPS Variable  3IOPS/GB
Min 100 IOPS
Max 6000 IOPS 
3IOPS/GB
Min 100 IOPS
Max 6000 IOPS

Storage percent


Sample threshold (percent or value): 80 percent

Friday, 11 January 2019

Performance troubleshooting using new Azure Database for PostgreSQL features

At Ignite 2018, Microsoft’s Azure Database for PostgreSQL announced the preview of Query Store (QS), Query Performance Insight (QPI), and Performance Recommendations (PR) to help ease performance troubleshooting, in response to customer feedback. This blog intends to inspire ideas on how you can use features that are currently available to troubleshoot some common scenarios.

This blog nicely categorizes the problem space into several areas and the common techniques to rule out possibilities to quickly get to the root cause. We would like to further expand on this with the help of these newly announced features (QS, QPI, and PR).

In order to use these features, you will need to enable data collection by setting pg_qs.query_capture_mode and pgms_wait_sampling.query_capture_mode to ALL.

Azure Certification, Azure Guides, Azure Tutorial and Material, Azure Learning

You can use Query Store for a wide variety of scenarios where you can enable data collection to help with troubleshooting these scenarios better. In this article, we will limit the scope to regressed queries scenario.

Regressed queries


One of the important scenarios that Query Store enables you to monitor is the regressed queries. By setting pg_qs.query_capture_mode to ALL, you get a history of your query performance over time. We can leverage this data to do simple or more complex comparisons based on your needs.

One of the challenges you face when generating a regressed query list is the selection of comparison period in which you baseline your query runtime statistics. There are a handful of factors to think about when selecting the comparison period:

◉ Seasonality: Does the workload or the query of your concern occur periodically rather than continuously?
◉ History: Is there enough historical data?
◉ Threshold: Are you comfortable with a flat percentage change threshold or do you require a more complex method to prove the statistical significance of the regression?

Now, let’s assume no seasonality in the workload and that the default seven days of history will be enough to evaluate a simple threshold of change to pick regressed queries. All you need to do is to pick a baseline start and end time, and a test start and end time to calculate the amount of regression for the metric you would like to track.

Looking at the past seven-day history, compared to last two hours of execution, below would give the top regressed queries in the order of descending percentage. Note that if the result set has negative values, it indicates an improvement from baseline to test period when it’s zero, it may either be unchanged or not executed during the baseline period.

create or replace function get_ordered_query_performance_changes(
baseline_interval_start int,
baseline_interval_type text,
current_interval_start int,
current_interval_type text)
returns table (
     query_id bigint,
     baseline_value numeric,
     current_value numeric,
     percent_change numeric
) as $$
with data_set as (
select query_id
, round(avg( case when start_time >= current_timestamp - ($1 || $2)::interval and start_time < current_timestamp - ($3 || $4)::interval then mean_time else 0 end )::numeric,2) as baseline_value
, round(avg( case when start_time >= current_timestamp - ($3 || $4)::interval then mean_time else 0 end )::numeric,2) as current_value
from query_store.qs_view where query_id != 0 and user_id != 10 group by query_id ) , 
query_regression_data as (
select *
, round(( case when baseline_value = 0 then 0 else (100*(current_value - baseline_value) / baseline_value) end )::numeric,2) as percent_change 
from data_set ) 
select * from query_regression_data order by percent_change desc;
$$
language 'sql';

If you create this function and execute the following, you will get the top regressed queries in the last two hours in descending order compared to their calculated baseline value over the last seven days up to two hours ago.

select * from get_ordered_query_performance_changes (7, 'days', 2, 'hours');

Azure Certification, Azure Guides, Azure Tutorial and Material, Azure Learning

The top changes are all good candidates to go after unless you do expect the kind of delta from your baseline period because, say, you know the data size would change or the volume of transactions would increase. Once you identified the query you would like to further investigate, the next step is to look further into query store data and see how the baseline statistics compare to the current period and collect additional clues.

create or replace function compare_baseline_to_current_by_query_id(baseline_interval_cutoff int,baseline_interval_type text,query_id bigint,percentile decimal default 1.00)
returns table(
     query_id bigint,
     period text,
     percentile numeric,
     total_time numeric,
     min_time numeric,
     max_time numeric,
     rows numeric,
     shared_blks_hit numeric,
     shared_blks_read numeric,
     shared_blks_dirtied numeric,
     shared_blks_written numeric,
     local_blks_hit numeric,
     local_blks_read numeric,
     local_blks_dirtied numeric,
     local_blks_written numeric,
     temp_blks_read numeric,
     temp_blks_written numeric,
     blk_read_time numeric,
     blk_write_time numeric
)
as $$

with data_set as
( select *
, ( case when start_time >= current_timestamp - ($1 || $2)::interval then 'current' else 'baseline' end ) as period
from query_store.qs_view where query_id = ( $3 )
)
select query_id
, period
, round((case when $4 <= 1 then 100 * $4 else $4 end)::numeric,2) as percentile
, round(percentile_cont($4) within group ( order by total_time asc)::numeric,2) as total_time
, round(percentile_cont($4) within group ( order by min_time asc)::numeric,2) as min_time
, round(percentile_cont($4) within group ( order by max_time asc)::numeric,2) as max_time
, round(percentile_cont($4) within group ( order by rows asc)::numeric,2) as rows
, round(percentile_cont($4) within group ( order by shared_blks_hit asc)::numeric,2) as shared_blks_hit
, round(percentile_cont($4) within group ( order by shared_blks_read asc)::numeric,2) as shared_blks_read
, round(percentile_cont($4) within group ( order by shared_blks_dirtied asc)::numeric,2) as shared_blks_dirtied
, round(percentile_cont($4) within group ( order by shared_blks_written asc)::numeric,2) as shared_blks_written
, round(percentile_cont($4) within group ( order by local_blks_hit asc)::numeric,2) as local_blks_hit
, round(percentile_cont($4) within group ( order by local_blks_read asc)::numeric,2) as local_blks_read
, round(percentile_cont($4) within group ( order by local_blks_dirtied asc)::numeric,2) as local_blks_dirtied
, round(percentile_cont($4) within group ( order by local_blks_written asc)::numeric,2) as local_blks_written
, round(percentile_cont($4) within group ( order by temp_blks_read asc)::numeric,2) as temp_blks_read
, round(percentile_cont($4) within group ( order by temp_blks_written asc)::numeric,2) as temp_blks_written
, round(percentile_cont($4) within group ( order by blk_read_time asc)::numeric,2) as blk_read_time
, round(percentile_cont($4) within group ( order by blk_write_time asc)::numeric,2) as blk_write_time
from data_set
group by 1, 2
order by 1, 2 asc;
$$
language 'sql';

Once you create the function, provide the query id you would like to investigate. The function will compare the aggregate values between the before and after based on the cutoff time you provide. For instance, the below statement would compare all points prior to two hours from now to points after the two hours mark up until now for the query. If you are aware of outliers that you want to exclude, you can use a percentile value.

select * from compare_baseline_to_current_by_query_id(30, 'minutes', 4271834468, 0.95);

If you don’t use any, the default value is 100 which does include all data points.

select * from compare_baseline_to_current_by_query_id(2, 'hours', 4271834468);

Azure Certification, Azure Guides, Azure Tutorial and Material, Azure Learning

If you rule out that there is not a significant data size change and the cache hit ratio is rather steady, you may also want to investigate any obvious wait event occurrence changes within the same period. As wait event types combine different wait types into buckets similar by nature, there is not a single prescription on how to analyze the data. However, a general comparison may give us ideas around the system state change.

create or replace function compare_baseline_to_current_by_wait_event (baseline_interval_start int,baseline_interval_type text,current_interval_start int,current_interval_type text)
returns table(
     wait_event text,
     baseline_count bigint,
     current_count bigint,
     current_to_baseline_factor double precision,
     percent_change numeric
)
as $$
with data_set as
( select event_type || ':' || event as wait_event
, sum( case when start_time >= current_timestamp - ($1 || $2)::interval and start_time < current_timestamp - ($3 || $4)::interval then 1 else 0 end ) as baseline_count
, sum( case when start_time >= current_timestamp - ($3 || $4)::interval then 1 else 0 end ) as current_count
, extract(epoch from ( $1 || $2 ) ::interval) / extract(epoch from ( $3 || $4 ) ::interval) as current_to_baseline_factor
from query_store.pgms_wait_sampling_view where query_id != 0
group by event_type || ':' || event
) ,
wait_event_data as
( select *
, round(( case when baseline_count = 0 then 0 else (100*((current_to_baseline_factor*current_count) - baseline_count) / baseline_count) end )::numeric,2) as percent_change
from data_set
)
select * from wait_event_data order by percent_change desc;
$$
language 'sql';

select * from compare_baseline_to_current_by_wait_event (7, 'days', 2, 'hours');

The above query will let you see some abnormal changes between the two periods. Note that event count here is taken as an approximation and the numbers should be taken within the context of the comparative load of the instance given the time.

Azure Certification, Azure Guides, Azure Tutorial and Material, Azure Learning

As you can see, with the available time series data in Query Store, your creativity is your limit to the kinds of analysis and algorithms you could implement here. We showed you some simple calculations by which you could apply straight forward techniques to identify candidates and improve. We hope that this could be your starting point and that you share with us what works, what doesn’t and how you take this to the next level.

Thursday, 10 January 2019

Streamlined development experience with Azure Blockchain Workbench 1.6.0

We’re happy to announce the release of Azure Blockchain Workbench 1.6.0. It includes new features such as application versioning, updated messaging, and streamlined smart contract development. You can deploy a new instance of Workbench through the Azure portal or upgrade existing deployments to 1.6.0 using our upgrade script.

Please note the breaking changes section, as the removal of the WorkbenchBase base class and the changes to the outbound messaging format will require modifications to your existing applications.

This update includes the following improvements:

Application versioning


One of the most popular feature requests from you all has been that you would like to have an easy way to manage and version your Workbench applications instead of having to manually change and update your applications as you are in the development process.

We’ve continued to improve the Workbench development story with support for application versioning with 1.6.0 via the web app as well as the REST API. You can upload new versions directly from the web application by clicking “Add version.” Note that if you have any changes in the application role name, the role assignment will not be carried over to the new version.

Azure Blockchain Workbench, Azure Guides, Azure Certification, Azure Tutorial and Materials

Azure Blockchain Workbench, Azure Guides, Azure Certification, Azure Tutorial and Materials

Azure Blockchain Workbench, Azure Guides, Azure Certification, Azure Tutorial and Materials

You can also view the application version history. To view and access older versions, select the application and click “version history” in the command bar. Note, that as of now by default older versions are read only. If you would like to interact with older versions, you can explicitly enable the previous versions.

Azure Blockchain Workbench, Azure Guides, Azure Certification, Azure Tutorial and Materials

Azure Blockchain Workbench, Azure Guides, Azure Certification, Azure Tutorial and Materials

New egress messaging API


Workbench provides many integration and extension points, including via a REST API and a messaging API. The REST API provides developers a way to integrate to blockchain applications. The messaging API is designed for system to system integrations.

In our previous release, we enabled more scenarios with a new input messaging API. In 1.6.0, we have implemented an enhanced and updated output messaging API which publishes blockchain events via Azure Event Grid and Azure Service Bus. This enables downstream consumers to take actions based on these events and messages such as, sending email notifications when there are updates on relevant contracts on the blockchain, or triggering events in existing enterprise resource planning (ERP) systems.

Azure Blockchain Workbench, Azure Guides, Azure Certification, Azure Tutorial and Materials

Here is an example of a contract information message with the new output messaging API. You’ll get the information about the block, a list of modifying transactions for the contract, as well as information about the contract itself such as contract ID and contract properties. You also get information on whether or not the contract was newly created or if a contract update occurred.

{
     "blockId": 123,
     "blockhash": "0x03a39411e25e25b47d0ec6433b73b488554a4a5f6b1a253e0ac8a200d13f70e3",
     "modifyingTransactions": [
         {
             "transactionId": 234,
             "transactionHash": "0x5c1fddea83bf19d719e52a935ec8620437a0a6bdaa00ecb7c3d852cf92e18bdd",
             "from": "0xd85e7262dd96f3b8a48a8aaf3dcdda90f60dadb1",
             "to": "0xf8559473b3c7197d59212b401f5a9f07b4299e29"
         },
         {
             "transactionId": 235,
             "transactionHash": "0xa4d9c95b581f299e41b8cc193dd742ef5a1d3a4ddf97bd11b80d123fec27506e",
             "from": "0xd85e7262dd96f3b8a48a8aaf3dcdda90f60dadb1",
             "to": "0xf8559473b3c7197d59212b401f5a9f07b4299e29"
         }
     ],
     "contractId": 111,
     "contractLedgerIdentifier": "0xf8559473b3c7197d59212b401f5a9f07b4299e29",
     "contractProperties": [
         {
             "workflowPropertyId": 1,
             "name": "State",
             "value": "0"
         },
         {
             "workflowPropertyId": 2,
             "name": "Description",
             "value": "1969 Dodge Charger"
         },
         {
             "workflowPropertyId": 3,
             "name": "AskingPrice",
             "value": "30000"
         },
         {
             "workflowPropertyId": 4,
             "name": "OfferPrice",
             "value": "0"
         },
         {
             "workflowPropertyId": 5,
             "name": "InstanceOwner",
             "value": "0x9a8DDaCa9B7488683A4d62d0817E965E8f248398"
         },
     ],
     "isNewContract": false,
     "connectionId": 1,
     "messageSchemaVersion": "1.0.0",
     "messageName": "ContractMessage",
     "additionalInformation": {}
}

WorkbenchBase class is no longer needed in contract code


For customers who have been using Workbench, you will know that there is a specific class that you need to include in your contract code, called WorkbenchBase. This class enabled Workbench to create and update your specified contract. When developing custom Workbench applications, you would also have to call functions defined in the WorkbenchBase class to notify Workbench that a contract had been created or updated.

With 1.6.0, this code serving the same purpose as WorkbenchBase will now be autogenerated for you when you upload your contract code. You will now have a more simplified experience when developing custom Workbench applications and will no longer have bugs or validation errors related to using WorkbenchBase. See our updated samples, which have WorkbenchBase removed.

This means that you no longer need to include the WorkbenchBase class nor any of the contract update and contract created functions defined in the class. To update your older Workbench applications to support this new version, you will need to change a few items in your contract code files:

◈ Remove the WorkbenchBase class.
◈ Remove calls to functions defined in the WorkbenchBase class (ContractCreated and ContractUpdated).

If you upload an application with WorkbenchBase included, you will get a validation error and will not be able to successfully upload until it is removed. For customers upgrading to 1.6.0 from an earlier version, your existing Workbench applications will be upgraded automatically for you. Once you start uploading new versions, they will need to be in the 1.6.0 format.

Get available updates directly from within Workbench


Whenever a Workbench update is released, we announce the updates via the Azure blog and post release notes in our GitHub. If you’re not actively monitoring these announcements, it can be difficult to figure out whether or not you are on the latest version of Workbench. You might be running into issues while developing which have already been fixed by our team with the latest release.

We have now added the capability to view information for the latest updates directly within the Workbench UI. If there is an update available, you will be able to view the changes available in the newest release and update directly from the UI.

Azure Blockchain Workbench, Azure Guides, Azure Certification, Azure Tutorial and Materials

Breaking changes in 1.6.0


◈ WorkbenchBase related code generation: Before 1.6.0, the WorkbenchBase class was needed because it defined events indicating creation and update of Blockchain Workbench contracts. With this change, you no longer need to include it in your contract code file, as Workbench will automatically generate the code for you. Note that contracts containing WorkbenchBase in the Solidity code will be rejected when uploaded.

◈ Updated outbound messaging API: Workbench has a messaging API for system to system integrations. We have had an outbound messaging API which has been redesigned. The new schema will impact the existing integration work you have done with the current messaging API. If you want to use the new messaging API you will need to update your integration specific code.

- The name of the service bus queues and topics has been changed in this release. Any code that points to the service bus will need to be updated to work with Workbench version 1.6.0.
- ingressQueue - the input queue on which request messages arrive.
- egressTopic - the output queue on which update and information messages are sent.
- The messages delivered in version 1.6.0 are in a different format. Existing code that interrogates the messages from the messaging API and takes action based on its content will need to be updated.
◈ Workbench application sample updates: All Workbench applications sample code are updated since we no longer need the WorkbenchBase class in contract code. If you are on an older version of Workbench and use the samples on GitHub, or vice versa, you will see errors. Upgrade to the latest version of Workbench if you want to use samples.

Wednesday, 9 January 2019

Teradata to Azure SQL Data Warehouse migration guide

With the increasing benefits of cloud-based data warehouses, there has been a surge in the number of customers migrating from their traditional on-premises data warehouses to the cloud. Microsoft Azure SQL Data Warehouse (SQL DW) offers the best price to performance when compared to its cloud-based data warehouse competitors. Teradata is a relational database management system and is one of the legacy on-premises systems that customers are looking to migrate from.

The Teradata to SQL DW migrations involve multiple steps. These steps include analyzing the existing workload, generating the relevant schema models, and performing the ETL operation. The intent of this discussed whitepaper is to provide guidance for these aforesaid migrations with emphasis on the migration workflow, the architecture, technical design considerations, and best practices.

Migration Phases


Azure SQL Data Warehouse, Azure Certification, Azure Learning, Azure Study Materials

The Teradata migration should pivot on the following six areas. Though recommended, proof of concept is an alternative step. With the benefit of Azure, you can quickly provision Azure SQL Data Warehouses for your development team to start business object migration before the data is migrated and speed up the migration process.

Phase one – Fact finding

Through a question and answers session you can define what your inputs and outputs are for the migration project.

Phase two – Defining success criteria for proof of concept (POC)

Taking the answers from phase one, you identify a workload for running a POC to validate the outputs required and run the following phases as a POC.

Phase three: Data layer mapping options

This phase is about mapping the data you have in Teradata to the data layout you will create in Azure SQL Data Warehouse. Some of the common scenarios are data type mapping, date and time format, and more.

Phase four – Data modeling

Once you’ve defined the data mappings, phase four concentrates on how to tune Azure SQL Data Warehouse. This provides the best performance for the data you will be landing into it.

Phase five: Identify migration paths

What is the path of least resistance? What is the quickest path given your cloud maturity? Phase five helps describe the options open to you and then for you to decide on the path you wish to take.

Phase six: Execution of migration

Migrating your Teradata data to SQL Data Warehouse involves a series of steps. These steps are executed in three logical stages, preparation, metadata migration, and data migration.

Migration solution


To ingest data, you need a basic cloud data warehouse setup for moving data from your on-premise solution to Azure SQL Data Warehouse, and to enable the development team to build Azure Analysis Cubes once the majority of the data is loaded.

Azure SQL Data Warehouse, Azure Certification, Azure Learning, Azure Study Materials

◈ Azure Data Factory Pipeline is used to ingest and move data through the store, prep, and train pipeline.
◈ Extract and load files via Polybase into the staging schema on Azure SQL DW.
◈ Transform data through staging, source (ODS), EDW and sematic schemas on Azure SQL DW.
◈ Azure Analysis services will be used as the sematic layer to serve thousands of end users and scale out Azure SQL DW concurrency.
◈ Build operational reports and analytical dashboards on top of Azure Analysis services to serve thousands of end users via Power BI.

This whitepaper is broken into sections which detail the migration phases, the preparation required for data migration including schema migration, migration of the business logic, the actual data migration approach, and testing strategy.

Sunday, 6 January 2019

SONiC: Global support and updates

SONiC (Software for Open Networking in the Cloud), our open switch OS, has been in the fast lane. A diverse group of community partners have actively engaged with us to contribute and support the evolvement of the software.

SONiC is considered a live organism, always evolving. Microsoft and the community is developing, refining, and making SONiC freely available to anyone running global scale or cloud-type networks or just have a healthy interest in advanced networking.

Being in control of the network fabric and particularly having a hardware agnostic approach across larger heterogenous networks is critical. SONiC was created to provide those foundational attributes we ourselves needed when we set out to build our global network which powers both Azure and our other cloud services.

Recently, SONiC has received several enhancements and updates, along with additions to the ecosystem contributing to SONiC’s success.

Let’s take a look at what is new.

Global support now available


We are excited to see SONiC and its sibling SAI (Switch Abstraction Interface) being adopted by many global network innovators. Recently, both Dell EMC and Mellanox announced that SONiC will feature as switch OS options for customers using their respective hardware and on top of this, they will both offer global support services for this new combination.

Microsoft Tutorial and Material, Microsoft Guides, Microsoft Learning, Azure Guides
“Dell EMC Networking has led the industry to an open networking paradigm, allowing customers to transform their IT operations on their own terms. Having SONiC as a validated option on our Open Networking hardware is a no-brainer as customers increasingly look to open source to gain even more flexibility and the scale to power massive cloud-based services such as Azure,” said Tom Burns, Senior Vice President and General Manager, Dell EMC Networking & Solutions.

With the introduction of enterprise-class support services, SONiC is moving quickly into mainstream territory. Users will now be able to take advantage of SONiC and its open source benefits, while at the same time enjoying high-quality global support. We look forward to welcome more contributors and maintainers offering similar services, as the SONiC community continues to grow.

Microsoft Tutorial and Material, Microsoft Guides, Microsoft Learning, Azure Guides
“Mellanox is a leading SONiC contributor offering years of experience in high speed networking, large system validation and Open Source. Our customers are looking for open systems based on Spectrum and Spectrum-2 that are easily tailored to the changing and demanding needs of the cloud. Mellanox is in the unique position to globally support advanced services from networking, scalable high-performance computing, advanced storage transport and Datacenter switching,” said Amir Preacher Executive Vice President Business Development, Mellanox Technologies.

Our decade long experience with running the global network powering both Azure and our other cloud services, have taught us many things. A key area crucial to operations at scale, remains to be monitoring and telemetry, which we have worked into and refined for SONiC over the years.

Pinpoint network issues with Everflow


Everflow is the brainchild of Microsoft researchers and networking engineers. It can be very difficult to diagnose tricky network issues such as packet drops, random latency spikes and network loops, and it’s not unusual for networking engineers to spend hours and even days working on pinpointing and resolving such problems.

Everflow is a packet-level telemetry system with unprecedented level of detail. Everflow adds a timestamp to packets at each stop and mirrors packets to a centralized collector for deep analysis. For Azure, Everflow is one of the most powerful tools we use to diagnose packet loss and latency issues in our global network.

Microsoft Tutorial and Material, Microsoft Guides, Microsoft Learning, Azure Guides

Figure 1. Everflow concept chart

No more resource overflows


A network switch has limited capacity to store all the various rules that make up the foundation of routing packets across a network. As the size of the networks increase, it is quite common to reconfigure the network to adopt more rules, routes, and nodes along the way. Through our own experience in building and running cloud scale networks, we have seen critical incidents caused simply by resource usage exceeding the limit of the ASIC, or brain of the switch.

To ease the pain and help address overflows, we have added Critical Resource Monitoring (CRM) to SONiC. It allows your switch to proactively flag and proactively issue alerts if strain on any resource is approaching its thresholds. Further, it enables a network engineer to proactively query the current state of any critical resource in the network. With CRM, critical configuration changes can be handled with safe taps on all resources.

Richer telemetry collection


Traditional network management systems are typically based on polling switch equipment to acquire telemetry via SNMP. This type of pull based telemetry is very inefficient, and hard to scale.

When running online services such as Azure, the dependency on network health and the ability to quickly resolve is obviously paramount. So, naturally we decided to build a more efficient way to monitor network states.

SONiC provides the data we need using streaming telemetry. In short, this means our hardware is proactively pushing real-time, structured, and analytics-ready data to the management system. Streaming telemetry greatly enriches the way we collect performance data. Its a natural fit to SONiC’s unique architecture by merely adding a containerized module with a centralized Redis database where telemetry data can be directly read out. The gRPC feature that streaming telemetry is based on was contributed to SONiC by community member Alibaba Group.

Microsoft Tutorial and Material, Microsoft Guides, Microsoft Learning, Azure Guides

Figure 2. SONiC streaming telemetry through gRPC

The features presented in this blog are all great examples demonstrating SONiC’s powerful ability to run diagnostics,  prevent network failures, and provide fast and flexible telemetry. As online services and application requirements become more and more sophisticated, the SONiC team at Microsoft and our highly committed developer community will continue to build, innovate, and refine the software, making the learnings we gather from building and operating  the most reliable network in public cloud available to you. Please stay tuned for future updates.

Thursday, 3 January 2019

Microsoft Azure portal December 2018 update

Here’s the list of December updates to the Azure portal:

Compute


Updated experience for virtual machine creation and management

This month brings a few updates that improve the usability of creating and managing virtual machines.

When creating virtual machines, you now have more flexibility for configuring virtual machine network parameters. We have revised the interface to give you more control over creating virtual networks, virtual network subnets, and address space.

Microsoft Azure, Azure Certifications, Azure Guides, Azure Learning, Azure Tutorial and Material

Configuring virtual network parameters for new VMs

We have also added the ability to specify the disk type of any new data disks during virtual machine creation.

Microsoft Azure, Azure Certifications, Azure Guides, Azure Learning, Azure Tutorial and Material

Specifying disk type

Finally, we redesigned the Disks overview page to look more like a standard Azure resource, with the most important information in the Essentials area at the top of the page.

We have also added charts for key disk metrics, including disk IOPS, throughput, and queue depth.

Microsoft Azure, Azure Certifications, Azure Guides, Azure Learning, Azure Tutorial and Material

New charts

To try out the new experience:

1. From the left-navigation menu, select Create a Resource.
2. Select any virtual machine image that you prefer.
3. On the Create a virtual machine page, select the Networking tab and then click Create New under the Virtual network name box.

Security


Security Center network map is now generally available

Security Center’s interactive network map provides a graphical view with security overlays giving you recommendations and insights for hardening your network resources. Using the map, you can see the network topology of your Azure workloads, connections between your virtual machines and subnets, and the capability to drill down from the map into specific resources and the recommendations for those resources.

Microsoft Azure, Azure Certifications, Azure Guides, Azure Learning, Azure Tutorial and Material

Network map

To open the Network map:

1. In the Security Center, under Resource Security Hygiene, select Networking.
2. Under Network map select See topology.

Updated Security Policy page

The Security Policy page was updated to reflect the built-in Azure Security Center policies as they are created in Azure policies. You can see the parameters for each of the policies that are assessed by Azure Security Center and configure existing security policies that apply to selected scopes (subscriptions or management groups).

Updated experience for Access control (IAM)

Let’s try this again. Controlling access to Azure resources using role-based access control (RBAC) is one of the most common tasks performed in Azure, and the experience for managing access is consistent across the Azure portal for different service types. We’ve updated the Access control (IAM) blade in the portal with a new interface based on tabs to improve performance and to help you complete important tasks such as checking a user's access more quickly. Here’s everything that’s changing in the IAM blade:

◈ Improved performance of the IAM blade
◈ A check access feature to quickly view role assignments for a single user, group, service principal, or managed identity
◈ Tiles that link to common tasks
◈ A deny assignments tab to view any relevant deny assignments. Deny assignments are read-only and can only be set by Azure.

Microsoft Azure, Azure Certifications, Azure Guides, Azure Learning, Azure Tutorial and Material

The new Access control (IAM) blade

To see the new IAM blade:

1. Select All services and select the scope or resource you want to view or manage. For example, you can select Management groups, Subscriptions, Resource groups, or any resource.

2. In the resource blade, select “Access control (IAM),” from the menu.

Management Tools


Updates to Azure Site Recovery

Azure Site Recovery now supports disaster recovery for Azure virtual machines deployed in Azure Availability zones. You can replicate zone pinned VMs from one Azure region to another region. If the target region supports Availability zones, you can configure the target VMs to be zone pinned VMs. If not, you can configure the VMs to be single instances or to be part of an availability set.

Microsoft Azure, Azure Certifications, Azure Guides, Azure Learning, Azure Tutorial and Material

Replicate your virtual machines to another Azure region.

To try out this feature:

1. Select any virtual machine deployed in an availability zone.
2. Select Disaster recovery in the left menu.
3. Review the defaults and select Enable replication.

Tuesday, 1 January 2019

Announcing general availability of Azure Machine Learning service: A look under the hood

Azure Machine Learning service contains many advanced capabilities designed to simplify and accelerate the process of building, training, and deploying machine learning models. Automated machine learning enables data scientists of all skill levels to identify suitable algorithms and hyperparameters faster. Support for popular open-source frameworks such as PyTorch, TensorFlow, and scikit-learn allow data scientists to use the tools of their choice. DevOps capabilities for machine learning further improve productivity by enabling experiment tracking and management of models deployed in the cloud and on the edge. All these capabilities can be accessed from any Python environment running anywhere, including data scientists’ workstations.

We built Azure Machine Learning service working closely with our customers, thousands of whom are using it every day to improve customer service, build better products and optimize their operations. Below are two such customer examples.

TAL, a 150-year-old leading life insurance company in Australia, is embracing AI to improve quality assurance and customer experience. Traditionally, TAL’s quality assurance team could only review a randomly selected 2-3 percent of cases. Using Azure Machine Learning service, it is now able to review 100 percent of cases.

“Azure Machine Learning regularly lets TAL’s data scientists deploy models within hours rather than weeks or months – delivering faster outcomes and the opportunity to roll out many more models than was previously possible. There is nothing on the market that matches Azure Machine Learning in this regard.”

Elastacloud, a London-based data science consultancy, uses Azure Machine Learning service to build and run the Elastacloud Energy BSUoS Forecast service, an AI-powered solution that helps alternative energy providers to better predict demand and reduce costs.

“With Azure Machine Learning, we support BSUoS Forecast with no virtual machines and nothing to manage. We built a highly automated service that hides its complexity inside serverless boxes.”

Azure Machine Learning service design principles


To simplify and accelerate machine learning, Azure Machine Learning has been built on the following design principles that are detailed in the rest of the blog.

◈ Enable data scientists to use a familiar and rich set of data science tools
◈ Simplify the use of popular machine learning and deep learning frameworks
◈ Accelerate time to value by offering end-to-end machine learning lifecycle capabilities

Familiar data science tools


Data scientists expect to use the full Python ecosystem of libraries and frameworks and the ability to train locally on their laptop or workstation. There are a wide variety of tools used across the industry, but broadly they fall under command line interfaces, editors and IDEs, and Notebooks. Azure Machine Learning service has been designed to support all of these. Its Python SDK is accessible from any Python environment, IDEs like Visual Studio Code (VS Code) or PyCharm, or Notebooks such as Jupyter and Azure Databricks. Let’s look deeper at the Azure Machine Learning service integration with a couple of these tools.

Jupyter Notebooks are a popular development environment for data scientists working in Python. Azure Machine Learning service provides robust support for both local and hosted notebooks (such as Azure Notebooks) and provides built-in widgets that allow data scientists to monitor the progress of training jobs visually in near real-time as shown in the image below. For customers who do machine learning in Azure Databricks, the Azure Databricks notebooks can be used just as well.

Azure Certification, Azure Tutorial and Material, Azure Guides, Azure Learning

Visual Studio Code is a lightweight but powerful source code editor which runs on your desktop and is available for Windows, macOS and Linux. The Python extension for Visual Studio Code, combines the power of Jupyter Notebooks with the power of Visual Studio Code. This allows data scientists to experiment incrementally in a “notebook style” while also getting all the productivity Visual Studio Code has to offer, such as IntelliSense, built-in debugger, and Live Share as shown in the image below.

Azure Certification, Azure Tutorial and Material, Azure Guides, Azure Learning

Support for popular frameworks


Frameworks are the most important libraries that a data scientist uses to build their models. Azure Machine Learning service supports all python-based frameworks. The most popular ones, scikit-learn, PyTorch, and TensorFlow, have been made into an Estimator class to simplify submission of training code to remote compute, whether it's on a single node or distributed training across GPU clusters. Furthermore, this is not just limited to machine learning frameworks. Any packages from the vast Python ecosystem can be used.

We realize that customers often face several challenges when they try to use multiple frameworks to build models and deploy them to a variety of hardware and OS platforms. This is because the frameworks have not been designed to be used interchangeably and require specific optimizations for hardware and OS platforms. To address these problems, Microsoft has worked with industry leaders such as Facebook and AWS, as well as hardware companies, to develop the Open Neural Network Exchange (ONNX) specification  for describing machine learning models in an open standard format. Azure Machine Learning service supports ONNX and enables customers to deploy, manage, and monitor ONNX models easily. Additionally, to  provide a consistent software platform to run ONNX models across cloud and edge, we announced that we will open source the ONNX runtime today. We welcome you to join the community and contribute to the ONNX project.

End-to-end machine learning lifecycle


Azure Machine Learning seamlessly integrates with Azure services to provide end-to-end capabilities for the machine learning lifecycle which include data preparation, experimentation, model training, model management, deployment, and monitoring.

Data preparation

Customers can use Azure’s rich data platform capabilities, such as Azure Databricks, to manage and prepare their data for machine learning. The DataPrep SDK is available as a companion to the Azure Machine Learning Python SDK to simplify data transformations.

Training

Azure Machine Learning service provides seamless distributed compute capabilities that allow data scientists to scale out training from their local laptop or workstation to the cloud. The compute is on-demand. Users only pay for compute time and don’t have to manage and maintain GPU and CPU clusters.

Azure Certification, Azure Tutorial and Material, Azure Guides, Azure Learning

Data professionals, who are already invested in Apache Spark, should train on Azure Databricks clusters. The Azure Machine Learning service SDK is integrated into the Azure Databricks environment and can seamlessly extend it for experimentation, model deployment, and management.

Experimentation

Data scientists create their model through a process of experimentation, iterating over their data and training code multiple times until they get the desired results from the model. Azure Machine Learning service provides powerful capabilities to improve the productivity of data scientists while also enhancing governance, repeatability, and collaboration during the model development process.

1. Using automated machine learning, data scientists can point to the dataset, and a scenario (regression, classification, or forecasting) and  automated machine learning uses advanced techniques to propose a new model by doing feature  engineering, selecting the algorithm and sweeping hyperparameters.

2. Hyper parameter-tuning of existing models enables fast and intelligent exploration of hyperparameters, with the early termination of non-performant training jobs, helps improve model accuracy.

3. Machine Learning pipelines allow data scientists to modularize their model training into discrete steps such as data movement, data transforms, feature extraction, training, and evaluation. Machine Learning pipelines are a mechanism for automating, sharing, and reproducing models. They also provide performance gains by caching intermediate outputs as the data scientist iterates in the model development inner loop.

Azure Certification, Azure Tutorial and Material, Azure Guides, Azure Learning

4. Finally, run history captures each training run, the model performance and the related metrics. We keep track of the code, compute, and datasets used in training the model. The data scientist can compare runs, and then select the “best” model for their problem statement. Once selected, the model is registered to the Model registry, which provides auditability – including provenance – of the models in production.

Deployment, model management, and monitoring


Once data scientists complete model development, there is the work of putting them into production, managing, and monitoring them. Azure Machine Learning service model registry keeps track of models and their version history, along with the lineage and artifacts of the model. 

Azure Machine Learning service provides the capability to deploy to both cloud and the edge, doing real-time and batch scoring depending on the customer’s needs. In the cloud, Azure Machine Learning service will provision, load balance, and scale a Kubernetes cluster using Azure Kubernetes Service (AKS) or attach to the customer’s own AKS cluster. This allows for multiple models to be deployed into production. The cluster will auto-scale with the load. Model management activities can be done with both the Python SDK, UX or with Command Line Interface (CLI) and REST API, which are callable from Azure DevOps. These capabilities fully integrate the model lifecycle with the rest of our customer’s app lifecycle.

Azure Certification, Azure Tutorial and Material, Azure Guides, Azure Learning

Models in the registry can also be deployed to edge devices with integration with Azure IoT Edge service.

Once the model is in production, the service collects both application and model telemetry that allows the model to be monitored in production for operational and model correctness. The data captured during inferencing is presented back to the data scientists and this information can be used to determine model performance, data drift, and model decay.

For extremely fast and low-cost inferencing, Azure Machine Learning service offers hardware accelerated models (in preview) that provide vision model acceleration through FPGAs. This capability is exclusive to Azure Machine Learning service and offers class-leading latency benefits, as well as cost/transaction benefits for offline processing jobs.