Unlocking ChatGPT’s Potential in IT Service Management

In today’s AI-driven landscape, ChatGPT has taken center stage, sparking discussions about its impact on jobs, its pros, cons, and the potential risks of unregulated AI. While debates abound, one cannot overlook the myriad opportunities ChatGPT offers in IT Service Management (ITSM). This blog aims to demystify ChatGPT, shedding light on how it works and exploring practical applications in ITSM. We’ll also discuss considerations for organizations looking to embrace this technology.

Understanding Generative AI and ChatGPT

Generative AI relies on algorithms to generate new data patterns from extensive training datasets. It includes various types such as Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), Recurrent Neural Networks (RNNs), and Transformer Models. ChatGPT falls under the category of Transformer Models, leveraging OpenAI’s Generative Pre-trained Transformer (GPT) architecture to produce context-aware text responses dynamically.

How ChatGPT Functions

ChatGPT operates by predicting word sequences in natural language to create grammatically correct and contextually relevant text. This process involves three key phases:

  1. Pre-training: ChatGPT learns language patterns from a vast corpus of internet text, without access to specific documents or confidential data. It predicts the next word in a sentence by discerning patterns in the data.
  2. Fine-tuning: The model undergoes fine-tuning by human reviewers following guidelines. This aligns its behavior with desired standards.
  3. Generating Responses: In response to input, ChatGPT generates a series of tokens (words) to form a response based on training patterns. It creates contextually relevant answers while paying attention to input words.

Benefits of ChatGPT

The benefits of ChatGPT in ITSM are akin to traditional AI and automation:

  • Enhanced User Experiences: ChatGPT-powered chatbots and virtual assistants elevate service quality for users and providers.
  • Workload Reduction: IT personnel experience reduced workloads, allowing them to focus on higher-value tasks and address skill shortages.
  • Improved Service Availability: Speedy issue resolution increases service availability, enhances productivity, and reduces downtime.
  • Cost Savings: Efficient resource allocation leads to tangible cost savings.

Practical ITSM Applications

Concrete ITSM use cases for ChatGPT include:

  1. Incident Management: ChatGPT provides automated initial support, guiding users through troubleshooting steps and resolving some issues without human intervention.
  2. Knowledge Management: Efficiently managing and searching an IT organization’s knowledge base, providing rapid and accurate solutions to IT staff and end-users.
  3. Service Request Management: Automation of tasks such as password resets and software installations, expediting request handling.
  4. Self-service Capabilities: Empowering chatbots to retrieve essential information from knowledge articles and diverse sources.
  5. IT Asset Management: Integration with IT asset tools for quick access to asset status, availability, and history.
  6. Automated Notifications: Prompt creation and dispatch of status updates on IT issues.
  7. Training: Generation of training materials and scenarios to equip IT teams for diverse challenges.

Considerations and Limitations

While ChatGPT offers immense potential, it’s important to be aware of its limitations:

  1. Data Limitations: ChatGPT’s responses are based on information available up to September 2021 and may not reflect recent events.
  2. Accuracy Challenges: It may occasionally generate incorrect or nonsensical answers and lacks ongoing contextual awareness.
  3. Quality of Training Data: Effectiveness depends on the quality and relevance of its training data.
  4. Inconsistencies: Responses may vary based on input phrasing, and the model might inadvertently reflect biases from its training data.
  5. Human Interaction: ChatGPT cannot replicate the nuances and empathy of human interaction, particularly in dealing with frustrated users.
  6. Response Length: Responses may become overly lengthy or complex, potentially overwhelming end-users.

In conclusion, ChatGPT’s integration into ITSM offers substantial benefits, but organizations should be mindful of its capabilities and limitations. By harnessing its power strategically, businesses can enhance user experiences and streamline IT operations.

Read more from IT Care Center here.


ITSM Maturity Assessment – The good, the bad and the ugly!

The intended aim of an ITSM Maturity Assessment is often very different from the outcome achieved.  Let me explain…..

Why do an assessment?

Organisations typically undertake assessments with the aim of achieving the following outcomes:

  • Identify improvement opportunities
  • Align IT with business goals
  • Enhance customer satisfaction
  • Optimise resource allocation
  • Increase operational efficiency
  • Reduce costs
  • Mitigate risks
  • Improve processes, tools and governance
  • Foster continual improvement

As you can see, ITSM Maturity Assessments are a great idea.

When done well, a good ITSM Maturity Assessment can easily achieve these outcomes.

In theory, they provide organisations with an opportunity to baseline their current ITSM capabilities, including processes, tooling, organisation design, people and governance.  In practice, they often fail to do deliver on their anticipated benefits and here’s why.

 

Method

Often the way in which organisations set about a process maturity assessment falls short due to several reasons.

Often, they’ll attempt to download a free self-assessment spreadsheet from the internet, which will rely upon objectivity from those undertaking the assessment.  Unfortunately, it’s just not in the best interests of our own teams to admit their shortcomings, so the results are often overly optimistic, even when undertaking by various team members with the results averaged.

Assessments and audits are often muddled.  An audit tends to lead to some form of certification or achievement of a standard.  It will rely upon a standard set of audit criteria which will not flex to the needs of those being audited.  An assessment, when done well, is flexible.  It adapts to the needs of the those being assessed and adjusts accordingly.  It is conversational, not based on a checklist.  And it isn’t conducted by an auditor.  It’s conducted by a consultant who understands that real-world service delivery excellence doesn’t have to “comply” with ITIL best practices or any other academic frameworks to be effective.

So, in essence, assessments are a good thing, but choose one which is impartial, objective, and flexible.

 

Challenges you will face

When undertaking an assessment, you will undoubtedly hit a number of challenges.

These might include:

  • Getting agreement to do the assessment
    As you will have read above, there are distinct advantages to an independent assessment, as opposed to a free self-assessment. However, there can be resistance to this.  Funding will certainly be an issue.  The cost versus benefit of the assessment must be sold to senior stakeholders.  In addition, the primary benefit of an assessment laying the foundations for a structured improvement programme must be emphasised.  It is important that all stakeholders understand that the assessment is step one of the improvement programme.
  • Setting the assessment scope
    Setting the correct scope of the maturity assessment is vital. As many of you will know, assessing ITSM maturity with just one or two process areas is impossible, since processes are intrinsically linked to one another. In my experience, a service lifecycle view must be adopted and the scope should reflect everything from demand management, project and service design, through transition and into live operation.  Often, upstream issues will be masked by issues early in the service lifecycle, such as shortcomings in service design disciplines.
    Scope must be carefully considered, not just from a process scope perspective, but also more broadly.  The presence of process diagrams is inconsequential if the culture, governance, tooling or skills have shortcomings.
    So we must consider a holistic approach when setting scope, which covers the right process areas through multiple lenses.
  • Making the time to undertake the assessment
    It is often difficult to see where to find the time to conduct the assessment, as they do require an investment in time from process owners, practitioners, IT leaders and in some cases, the team’s customers / end users. However, the long-term view must be considered here.  Whilst there will be a short term hit on productivity, the assessment will almost certainly provide insights as to where improvements and efficiency gains can be found, that will more than repay than initial investment in time.
  • Reaching consensus on the results
    Once the assessment is over, the results will be published. Occasionally, there will be criticism which will be hard to hear.  Like being told that your baby is ugly, negative feedback is seldom welcome.  However, it is necessary not only to embrace all the feedback, but to reach consensus upon on it.  Only at this time can the team unite to embrace the recommendations and embark upon an improvement programme.
  • Making progress on the recommendations
    There is a risk that the assessment and related recommendations will sit on a shelf after the assessment and the effort involved will be wasted. It is critical that responsibility is assigned to managing and tracking the implementation of the recommendations.  We’d recommend setting up a governance committee to ensure that progress is made, with a member of senior IT leadership as sponsor.

 

Output

Once the assessment is complete, what results do you get?  A 1 to 5 score with no commentary is worthless.  Results need to provide a balanced presentation the positives and the improvement areas.  A maturity score is great, but only if the marking basis is consistent.

The output shouldn’t just contain a list of recommendations.

Recommendations need to be prioritised, preferably by impact, urgency and effort.  Ideally, they should be presented in a roadmap format, to enable a clear improvement pathway to be consumed visually. A roadmap diagram is a great way of showing the journey ahead and also of tracking progress.

The recommendations and roadmap diagram can then form the basis of an implementation plan.  This plan becomes the focal point of the improvement project.

Assessment recommendations should be:

  • Assigned to an individual owner
  • Tracked on a regular basis
  • Overseen by an appropriate governance committee
  • Supported by expert advice and guidance to enable them to be implemented
  • Support by evidence to support the assertion that they have been completed.

 

Tracking progress

Progress can be tracked in a number of ways.

  • Progress against individual recommendations
  • Progress against groups of recommendations relating to a specific process area (e.g. Change Management)
  • Re-assessment at regular intervals to refocus continual improvement efforts

You can read more about Syniad’s assessment services on their web site, or contact them for more information.


Rescue ROI: see how GoTo customers achieved 395% ROI in 3 years with payback in 6 months

Here at GoTo, we see the value that LogMeIn Rescue, our secure, enterprise-grade remote support solution, brings to our customers every day. But we wanted to know, what’s the actual return on investment (ROI) our customers have achieved?  

We had a hunch it would be good, but we were surprised when we first saw the numbers from our 2023 Total Economic ImpactTM (TEI) study that we commissioned from Forrester Consulting. So surprised in fact, we asked to run the numbers twice.  

By using Rescue remote support, the study found a three-year 395% ROI with payback in less than six months. 

Why is ROI important in technology investments? 

When it comes to investing in technology, it’s important to consider the return on investment. After all, you want to make sure that the money you’re putting into technology is actually going to pay off in the long run. Of course, it’s important to keep in mind that ROI isn’t just about the money. It can also be about time saved or efficiency gained. 

One way to ensure that you’re getting a good ROI is by doing your research beforehand. Look at the technology you’re considering and see how it’s worked out for other businesses. Have they seen a significant increase in productivity or revenue since implementing it? If so, it’s likely a good investment.  

A closer look at Rescue’s ROI  

The results from Forrester’s study are based on a composite organization created from aggregated interview and survey responses. The case study organization is a $1.5 billion multinational enterprise with 5,000 employees, including 50 internal and 60 customer support technicians. 

Key takeaways: 

Faster resolutions lead to greater productivity 

By using Rescue, on average, surveyed organizations were able to resolve computer problems 57% faster and mobile problems 23% faster.  

Lost productivity time for computer issues between queue time and resolution time was cut in half. For mobile problems, lost productivity time was cut by a quarter.  

Shifting left saves time and money 

By using Rescue to resolve issues remotely, the number of on-site visits fell by 15%.  

Additionally, since Rescue enabled remote technicians to go further in their troubleshooting, often finding the cause of the issue more accurately, it took on-site technicians less time to repair the problem, resulting in 45% cost savings.  

Customer satisfaction and loyalty strengthens the bottom line 

Rescue reduced the average time to restore customers’ computer issues by 35% and mobile issues by 25%.  

These impressive results led to a 21% customer satisfaction (CSAT) score increase, and a 28% improvement in Net Promoter Scores (NPS)1 

A 26% increase in Customer Effort Scores (CES) demonstrates the positive impact of the product, as customers don’t need to struggle to load a software product for technicians to connect to their device and resolve their issue.  

View the infographic here.


Beyond the Basics: Unconventional Applications of ITSM Solutions

IT service management (ITSM) solutions have long been associated with streamlining IT operations, incident management, and IT-related processes. However, these versatile tools have the potential to be adapted for a range of other applications that extend well beyond their traditional use cases.

Some of these uses are well understood and make sense, while others may be less explored. Getting beyond the basics of your ITSM solutions is easier than you might imagine. The biggest challenge may be overcoming cultural barriers within your organization rather than trying to determine new use cases.

This article explores creative and lesser-known ways organizations can leverage ITSM and ESM solutions to enhance efficiency and effectiveness across various departments.

Facilities Management

Facilities management plays a pivotal role in ensuring the smooth operation of any organization. By harnessing the capabilities of ITSM solutions, businesses can revolutionize their facilities management processes.

These solutions provide a structured framework for managing maintenance requests, tracking space allocation, inventorying equipment, and efficiently reserving rooms. The centralization of these tasks streamlines communication among teams, enhances transparency in resource allocation, and optimizes the utilization of physical assets.

As a result, organizations can reduce downtime, minimize disruptions, and enhance overall operational efficiency.

HR Onboarding and Offboarding

The employee lifecycle, from onboarding to offboarding, involves numerous intricate tasks that can benefit from the automation and organization provided by ITSM solutions. By integrating ITSM workflows into HR processes, organizations can facilitate a seamless transition for new and departing employees.

These workflows can automate the provisioning of system access, streamline equipment setup, and facilitate the management of essential documentation. The result is a standardized and efficient process that minimizes manual errors, accelerates time-to-productivity for new hires, and ensures the secure and compliant offboarding of departing employees.

Vendor and Supplier Management

Effective vendor and supplier management is vital for maintaining a reliable supply chain. While typically associated with IT procurement, ITSM solutions can be adapted to handle vendor relationships across various departments.

By leveraging these tools, organizations can centralize vendor information, track contracts, monitor adherence to service level agreements (SLAs), and measure supplier performance. This structured approach fosters transparency, enhances vendor collaboration, and helps organizations make informed decisions about their partnerships.

Project Management

The principles of ITSM can be seamlessly applied to project management, extending the benefits of structured processes and automation to cross-functional projects. Organizations can ensure that projects are executed efficiently and on schedule by configuring ITSM workflows for project initiation, task assignment, resource allocation, and milestone tracking.

Collaboration is enhanced as team members follow standardized processes, improving project outcomes and increasing stakeholder satisfaction.

Knowledge Base for Customers

ITSM solutions excel in knowledge management, and this capability can be harnessed to create powerful customer support resources. By utilizing these tools, organizations can build comprehensive knowledge bases accessible to customers.

This empowers users to find solutions independently, reducing the strain on support teams and offering quick resolutions to common inquiries. The result is enhanced customer satisfaction, reduced support costs, and the cultivation of a self-reliant customer base.

Marketing Campaign Tracking

While not an obvious application, ITSM solutions can effectively manage marketing campaigns by providing structure and visibility to the entire process. By creating ITSM workflows, organizations can track campaign tasks, approvals, deadlines, and assets.

This approach facilitates collaboration among marketing teams, ensures the timely execution of campaigns, and allows for real-time progress monitoring. As a result, organizations can improve campaign management, enhance marketing strategies, and achieve better campaign outcomes.

Compliance and Security Audits

Maintaining compliance with regulations and security standards is a perpetual challenge for organizations across industries. ITSM solutions can serve as powerful allies in this endeavor by facilitating structured workflows for compliance checks and security audits.

By integrating these processes into the ITSM framework, organizations can systematically assess and address compliance gaps, thus reducing the risk of non-compliance and ensuring that security measures are consistently upheld.

Event Management

Whether internal team gatherings or external conferences, events demand meticulous planning and execution. ITSM solutions can streamline event management by centralizing tasks such as event logistics, registration tracking, session scheduling, and attendee management. These tools enable efficient coordination among event organizers, enhance attendee experiences, and provide a post-event analysis and improvement platform.

Internal Request Management

Beyond IT-related matters, ITSM solutions can revolutionize the management of various internal requests. From booking meeting rooms and requesting office supplies to submitting maintenance tickets, these solutions offer a standardized approach to request handling. Organizations can optimize resource allocation, reduce administrative overhead, and enhance overall operational efficiency by automating request routing, approval workflows, and notifications.

Telecom Expense Management

Telecom expenses can quickly spiral out of control without proper management. ITSM solutions can be adapted to track telecom usage, manage invoices, and optimize telecom plans. By implementing these workflows, organizations can better see telecom expenses, identify cost-saving opportunities, and ensure that resources are allocated effectively.

Conclusion

The realm of ITSM solutions extends far beyond their conventional applications. Organizations can harness their versatile workflows and automation capabilities to transform operations across departments.

ITSM solutions provide a robust framework for enhancing efficiency, standardizing processes, and optimizing resource allocation from facilities and HR to vendor management and beyond. By thinking outside the box and customizing these solutions to fit unique organizational needs, businesses can unlock new avenues of operational excellence and drive lasting success.

By Ruben Franzen, president, TOPdesk US


Meet the Brand: Efecte

First time exhibitors to SITS 23, Efecte help people to digitalise and automate their work. Customers across Europe leverage their cloud service to operate with greater agility, to improve the experience of end-users, and to save costs.

The use cases for their solutions range from IT service management and ticketing to improving employee experiences, business workflows, and customer service. Efecte are dominating the European market, with their headquarters in Finland, and regional hubs in Germany, Poland and Sweden.

How did your story with SITS start?

We were looking for the main ITSM event in the UK, and Marxtar with previous experience of SITS recommended them as the largest and best event to be part of in the UK. So SITS was our first choice to showcase Efecte.

As Efecte enter the UK market, we understood that to be part of the main ITSM event would raise Efecte’s profile and increase brand awareness and would also give us a chance to be present alongside other ITSM solution providers.

 What has been the value of exhibiting at SITS?

We believe that the Efecte brand is better known in the UK now, and we have gained some good leads from the event that we have been speaking to, we are in the process of turning them into customers.

How has exhibiting at SITS helped you elevate your brand?

It's hard to measure, but we believe more people in the UK now know Efecte as a service management solution provider.

Meet the Efecte team at SITS 2024 on 17-18th April at ExCeL London.


Why DevSecOps and what’s different about it? Security is not a ‘consideration’

Aiming for a faster, higher-quality, software development lifecycle (SDLC), DevOps has become the mainstream approach in recent years. Utilising Agile methodologies, development and operations teams collaborate throughout the entire process of developing, deploying, and managing applications. Alongside the growth of DevOps, there’s an increase in cloud migration, sophisticated cloud-native infrastructures and using a microservices approach with organisations eagerly adopting containerisation and kubernetes. The very nature of the new SDLC approach and these advances means security is not a ‘consideration’; it cannot be the ‘add on’ or afterthought. It is far more than that.

Here are three glaring examples of why DevSecOps - security as a central part of the entire lifecycle – is essential:

  • With hackers always on the lookout for the opportunity to penetrate code and DevOps faster cycle of code releases, embedding of security principles and practices must be in place at the very beginning of the lifecycle, when an application or solution is being planned. Rather than relying solely on testing and a security audit close to the release stage, developers must also be responsible for thinking about security.
  • With much of the cloud-native infrastructures having less defined network boundaries and offering a wider attack surface for cyber threats, it makes sense that investment of time and resources into security happens at each stage of the lifecycle, when issues are still easier, faster, and less expensive to fix, rather than to fix them retrospectively much later, right before production.
  • With increased collaboration between teams as part of a DevOps culture, this means new levels of sharing information are required whether its API tokens, access credentials or SSH keys. Keeping data secure becomes increasingly demanding and a new approach is needed to avoid attackers or carelessness causing serious damage.


Embracing AI in mental health: amplify engagement while balancing human touch!

Today, I’m looking at how AI’s making waves in the mental health space.

Exciting, scary, intriguing – all of the above!

Now, I know what you’re thinking, “Nick, isn’t AI all about robots and sci-fi stuff?” Well, yes and no. It’s incredible how smart these machines can be, picking up on patterns we humans might miss. They can analyse all sorts of data from social media, health records, and surveys, often providing early warning signs of mental health struggles. I’ve been amazed by the conversations I am having in this space with technology professionals!

Some AI tools are now providing on-the-go mental health support. We’re talking mood-tracking apps, mental health chatbots, even virtual reality therapy! But hold your horses! These aren’t or shouldn’t (in my opinion!) be replacing our good old therapy sessions – they’re just giving us some extra tools to help manage our mental health.

But let’s get one thing straight here – AI, while smart, is no substitute for a real-life, breathing professional. If you’re facing a tough time or a crisis, reach out to a professional, a trusted friend, a loved one. The point here is that AI can be your ally, your helping hand, but it should never be your only port of call.

I guess that’s my message today…

So, here are your quick-fire takeaways I want to leave you with – some food for thought:

1. AI can help spot early warning signs by analysing loads of data.
2. AI tools like apps and chatbots can help manage our mental health day-to-day.
3. AI’s a nifty tool, but it’s no replacement for professional advice.

The moral of the story is this – don’t shy away from using AI as part of your mental health toolkit. But do it wisely, thoughtfully, and never, ever forget the value of reaching out to a professional.

Just add it to your ‘playbook’ – those who know, know! 😂

Remember, we’ve got a load of tools at our disposal – it’s about using the right one at the right time.


Getting the best IT suppliers and getting the best from them – how SIAM is gaining momentum

Most organizations use multiple suppliers as part of their IT development and delivery, gaining access to best of breed skills and competitive pricing. But additional complexity also brings challenges and management overhead. Service integration and management (SIAM) is a management approach that helps organizations build an operating model that delivers maximum value from a diverse supply chain. Extending service management thinking outside organizational boundaries, SIAM is widely adopted in both public and private sectors around the world.

According to the 2022 Global SIAM survey, the top 5 strategic drivers for SIAM adoption are:

  • Better ability to measure and attribute service quality
  • Wanting to have better performance from existing vendors
  • Wanting to have more control of existing vendors
  • Better ability to measure and attribute service costs
  • Moving from a single vendor to multiple vendors

The top 5 benefits that organizations report having achieved include:

  • Better reporting and management information
  • Better collaboration between suppliers
  • Better supplier performance
  • Easier to add and remove suppliers
  • Spending less time on general supplier management

Scopism is the organization at the heart of the SIAM community. If you’d like to learn more about SIAM, Scopism offers many resources, including the conference and online community we’ll discuss in this article.

As organizations adopt digital technologies and strategies based on digital transformation, Scopism is seeing more and more interest in SIAM and how to incorporate service integration thinking into an IT operating model.

For a more detailed introduction to SIAM, you can read this blog or watch this short video.

ServiceNorth Global SIAM Conference

Taking place in Manchester, UK on November 7th, ServiceNorth is the world’s leading SIAM conference. Now in its sixth year, the event welcomes SIAM practitioners from all over Europe and beyond, and provides an engaging way to build your network, learn from speakers and meet leading SIAM organizations in the sponsor exhibition. This year’s agenda includes:

  • HM Land Registry SIAM case study
  • An agile approach to SIAM procurement
  • SIAM simulation introduction
  • Global SIAM survey panel discussion
  • And much more

Our sponsors include HCL Software, Infosys, 2Grips, Syniad IT, Digital Clarity and our platinum sponsor 4me. Learn more about the event and book your ticket here.

If you can’t attend in person, live stream tickets are also available.

Online SIAM Community

Scopism has also launched an online SIAM community in 2023. The community is free to join thanks to our community partners including HCL Software and Infosys – you can register as a member here.

The community includes:

  • Access to free downloads such as the SIAM Body of Knowledge and whitepapers
  • Online live streaming events including Ask the Expert and SIAM+ deep dives
  • ‘Ask the Community’: an opportunity to ask and answer SIAM questions
  • Regional and special interest groups
  • Videos from previous conferences
  • Create your profile and network with other SIAM practitioners

Author: Claire Agutter is the founder of Scopism and a service management speaker, author and consultant. In 2018-23 she was nominated by Computer Weekly as one of the most influential women in tech.


Expert insight: Why DevSecOps and what’s different about it? Security is not a ‘consideration’

Aiming for a faster, higher-quality, software development lifecycle (SDLC), DevOps has become the mainstream approach in recent years. Utilising Agile methodologies, development and operations teams collaborate throughout the entire process of developing, deploying, and managing applications. Alongside the growth of DevOps, there’s an increase in cloud migration, sophisticated cloud-native infrastructures and using a microservices approach with organisations eagerly adopting containerisation and kubernetes. The very nature of the new SDLC approach and these advances means security is not a ‘consideration’; it cannot be the ‘add on’ or afterthought. It is far more than that.

Here are three glaring examples of why DevSecOps - security as a central part of the entire lifecycle – is essential:

  • With hackers always on the lookout for the opportunity to penetrate code and DevOps faster cycle of code releases, embedding of security principles and practices must be in place at the very beginning of the lifecycle, when an application or solution is being planned. Rather than relying solely on testing and a security audit close to the release stage, developers must also be responsible for thinking about security.
  • With much of the cloud-native infrastructures having less defined network boundaries and offering a wider attack surface for cyber threats, it makes sense that investment of time and resources into security happens at each stage of the lifecycle, when issues are still easier, faster, and less expensive to fix, rather than to fix them retrospectively much later, right before production.
  • With increased collaboration between teams as part of a DevOps culture, this means new levels of sharing information are required whether its API tokens, access credentials or SSH keys. Keeping data secure becomes increasingly demanding and a new approach is needed to avoid attackers or carelessness causing serious damage.

The three key failure metrics your service desk needs to start tracking

Metrics are at the heart of IT service management, delivering insights on operations and helping identify areas of continual improvement. The usual service desk metrics help showcase the internal operational efficiency. For example, SLA, that measures the number of tickets resolved under the specified time is a key factor that showcases service desk efficiency. On the other hand, failure metrics help teams identify weak chinks in the IT infrastructure and help evaluate responses to failure events. This helps IT teams minimize the cascading effect that failures can cause on critical systems.

What are the key failure metrics to be tracked? In this article we will see the following three KPIs:

  • Mean time between failure

  • Mean time to failure

  • Mean time to repair

Mean time between failure (MTBF)

When there are frequent failures on IT infrastructure assets, be it networks, servers, workstations, etc., they have a cascading impact on the availability of IT and business services. These disruptions lead to loss of revenue and reputation. If a particular IT asset sees frequent downtimes, repair or replacement is often required. Before that, it helps to investigate and understand why the asset goes down often and in what circumstances. This helps plan asset maintenance and improve systems availability. MTBF is the metric that helps identify downtime causes and helps mitigate them or plan for quick recovery and better availability of IT systems.

Figure 1. Mean time between failure

If the MTBF of a particular IT asset is low, it means the asset sees frequent downtimes leading to IT and business disruptions.

MTBF example

In an organization, new updates to the storage drive kept failing whenever new Windows firmware updates were applied. This occurred a few times and the MTBF became worse. After analysing the issue, the team determined that the third-party driver caused the API required to carry out the update to either not be implemented, or to be faulty. When a new update is scheduled, if third-party drivers do not implement the necessary APIs, there are two possible solutions to explore. Swapping the APIs with the Windows alternatives for SATA and NVMe storage protocols, or obtaining a new and better supported version of the driver from the OEM can help implement updates, fix bugs, and close security loopholes. Monitoring and tracking driver upgrades and downtime helps improve the availability of the storage drives.

How to improve MTBF

  1. Implement a process to observe asset health to track and monitor failures. This helps identify the cause of disruptions.

  1. Analyze the root cause of the problem to create awareness, address long-term causes, and improve asset performance.

  1. Create a quick response strategy to effectively tackle and reduce downtimes that impact operations. The objective is to achieve fewer and more time between disruptions.

Mean time to failure (MTTF)

Assets failing regularly can interrupt your organization's IT operations, and result in the deterioration and underperformance of IT infrastructure. The MTTF metric helps determine the typical lifespan of an asset, device, or component. For IT assets and components with a low MTTF, it is often more time-efficient, and minimizes operational impacts and costs, to replace the IT component instead of fixing the component.

This applies especially to IT components linked to crucial operational elements of the infrastructure like a mainframe server stack or a network access point.

Figure 2. Mean time to failure

If the MTTF of an asset is unfavorable and fails regularly, it indicates that the IT asset is unreliable and needs frequent replacement to avoid impacting IT operations.

MTTF example

In an IT software development company, when a cable was connected or disconnected from the switch in the data and network server stack, the network cables would get loose, and disconnect or get damaged. This led to files becoming corrupted due to interrupted data transfer. Further analysis by the network team revealed that the snagless plastic cover kept breaking on the CAT6 RJ45 patch cable. This was due to the cable being procured from a manufacturer who used cheap material. The IT team then replaced the old cables with cables of better quality to make sure there would be no issues, like the loss or corruption of data, in future when cables are moved. This is a classic example, but tracking the MTTF of the cable on a regular basis helps IT teams understand the impact of critical assets, like components, so they can make informed decisions about repair and replacement.

How to increase MTTF

  1. Increase the asset life span by procuring assets of high quality and decommissioning assets of low quality and cost.

  1. Prevent large-scale disruptions to business operations by scheduling regular checks on components linked to critical assets.

  1. Implement a just-in-time inventory process that estimates the time an asset is operational, leading to reduced overhead costs for asset storage.

Mean time to repair (MTTR)

When a critical IT system fails, IT teams must get the system running as soon as possible. Delays in restoring IT systems can lead to loss of revenue and impact critical business operations. A well-organized recovery and response system can help IT teams respond to unplanned downtime and restore operations effectively. MTTR measures the average time taken to repair or troubleshoot an asset and return it to its operational capability.

Figure 3. Mean time to repair

The cost of a downtime increases as the MTTR increases. High MTTR suggests that your recovery and response operations are not quick and effective. System failures are unavoidable, but MTTR enables teams to react to asset failures in a timely and strategic way.

MTTR example

A software company faced a zero-day attack on a video game it was developing due to vulnerability in a code. The attack disrupted operations like Wi-Fi and surveillance systems. This led to the attackers accessing the organizations' network domain and confidential business files. The cybersecurity team informed employees about zero-day attacks and where they could report them. Every IT asset in the organization was equipped with next-generation antivirus (NGAV). The attack disabled the LAN and employee self-service portal, crippling the operations of the organization. Within an hour of the attack, the cybersecurity team was informed and helped by NGAV's ability, which leverages threat analytics and behavior patterns of users, and identified the suspicious activity. The cybersecurity team immediately ran a patch management script to rectify the vulnerability in the code, and locked down its on-premises network to avoid further impact operations and data theft.

How to reduce MTTR

  1. An efficient asset management strategy helps drive better decision-making by identifying bottlenecks, and designating that assets be repaired or replaced. This saves money and storage space.

  1. Define the responsibilities and roles for technicians to streamline the incident detection and resolution process.

  1. Provide technicians with detailed standard operating procedures to reduce miscommunication and confusion during a downtime.

  1. Measure MTTR using an Enterprise Asset Management solution that centralizes asset maintenance and monitoring information. This also helps optimize the utilization of assets, collect asset data, and predict possible downtime.

Conclusion
These failure metrics help teams identify the bottlenecks in operations and their responsiveness to incidents. They empower IT teams to achieve higher operational efficiency by pinpointing the root cause of persistent incidents. IT teams can improve their incident response strategy with a clear picture of areas where IT operations are impacted. These metrics can be implemented in organizations by using them as KPIs rather than just performance objectives. The metrics point out areas for process simplification and operational improvements, and are not merely targets to hit.

A quick summary of each metric:

  • MTBF provides better insights into your service desk's effectiveness at preventing future disruptions.

  • MTTF helps you understand the lifecycle of an asset and its reliability.

  • MTTR indicates the time spent on repairing and how quickly your IT teams are able to diagnose disruptions.

This article was originally published on servicedeskplus.com, on April 17, 2023, written by Saket Pasumarthy.