One of the things I try to challenge my teams with is following through on issues and user inquiries. There are many times when issues come our way, just to find out that it's really within the scope of another team to correct or address it. This can happen for several reasons. For example, from experience, a business user has determined that your team provides the best turnaround time on issues. It could also be that the documentation on whom to contact might be unclear and the business users goes knocking on the first door she finds.
In many cases, I see teams simply forward the e-mail or ticket along to another group. A lot of times, the e-mail or ticket won't have full documentation on timeline, impact, history, etc. The team receiving the e-mail or ticket might not react with the right level of urgency. In fact, I've seen issues go on for days like this; being passed from one team to another. In the meantime the business user just waits in frustration.
A better approach to prevent long running e-mail threads that lead nowhere, is for the receiving team to follow through on the issue, as if it were their own. In my opinion, if the business user sends an issue your way, you should own it to completion. Instead of forwarding the e-mail or ticket, get the right team(s) on a call, and ask the right questions. Communicate the right level of urgency on that call as well. Just choosing the right forum (phone call versus e-mail) can help significantly cut down on the turnaround time.
Once you have an answer, personally deliver it to the business user. Don't expect other teams to do it. They might not have the same finesse and level of service that you have. Remember, you always want to keep those business users delighted.
Many teams will complain that they don't have enough staffing to personally handle each of these types of issues or inquiries. I would challenge that assumption. Many times, we spent more time fighting fires and explaining bad results than it would have taken to just manage the issue to completion. Guess who the business user will complain about if the issue doesn't get addressed on time?
So skip that coffee or tea break if you have to. Challenge yourself to provide your users the best service possible. They'll thank you for it and your organizational growth will, indeed, reflect it (so will your bottom line).
Are you passionate about making your Production Support team better? Join me as we explore topics in Production Support of Mission Critical applications.
Wednesday, September 25, 2013
Tuesday, September 24, 2013
Are You Sure About Your Monitoring?
Today we had an embarrassing issue happen. It started at 1:00 AM and we didn't catch the problem until 10:00 AM when business users reported they were missing some data. So, basically, we went about half a day without knowing something was wrong. As it turns out, we had a monitoring gap.
A log monitor which captures certain strings in the file did not capture one of the strings it was configured for. Here's the timeline of events of why it didn't capture the error:
Gotcha! Clearly we missed this in our thinking when we set up the monitor. We've now configured our monitoring to always look at the last two log files. Since the files don't grow too quickly, that should suffice (given the 5 minute interval).
So if you use Sitescope, keep this in mind. Don't get caught with your pants down. I'm sure by now I've lost everyone who uses Sitescope for monitoring (they're now checking they don't have similar gaps). :-)
Cheers!
A log monitor which captures certain strings in the file did not capture one of the strings it was configured for. Here's the timeline of events of why it didn't capture the error:
- The monitor was set up to tail the log file every 5 minutes to capture everything in the log since the last time the monitor ran. This is by design with a vended application we use for monitoring.
- The monitor ran at 12:59 PM and didn't find any errors.
- The error comes in at 1:00:59 AM with the string "LOG EXCEPTION"
- The log file rolls because it has a size limitation.
- The monitor runs at 1:04 AM and tails the file again, but the error is now in the rolled file.
- 10:00 AM, the business user reports the problem
Gotcha! Clearly we missed this in our thinking when we set up the monitor. We've now configured our monitoring to always look at the last two log files. Since the files don't grow too quickly, that should suffice (given the 5 minute interval).
So if you use Sitescope, keep this in mind. Don't get caught with your pants down. I'm sure by now I've lost everyone who uses Sitescope for monitoring (they're now checking they don't have similar gaps). :-)
Cheers!
Friday, September 20, 2013
A Tough Interview Question and How to Ace It
I was asked the following question in a forum, and I thought it was worth sharing as a blog post:
Question (edited):
During a job interview I was given the following scenario to test my ability at handling difficult situations as a Production Support analyst.
You receive a call from two business users and:
Issue 1: The business user is saying they are unable to log into the application. This is happening to multiple business users and you know this is occurring during peak hours.
Issue 2: After the first call, you get a call from a different business user saying they are unable to generate reports to validate the data for another application.
Which issue should be given priority? How would you handle this situation?
Answer:
Several things are going on here:
I hope you find this answer helpful and that it will help you ace your next Production Support job interview!
Question (edited):
During a job interview I was given the following scenario to test my ability at handling difficult situations as a Production Support analyst.
You receive a call from two business users and:
- You are alone covering the shift.
- The issues are not documented in the run book.
- Both the stakeholders are insisting their issue is critical.
- The severity level of both issues is same.
Issue 1: The business user is saying they are unable to log into the application. This is happening to multiple business users and you know this is occurring during peak hours.
Issue 2: After the first call, you get a call from a different business user saying they are unable to generate reports to validate the data for another application.
Which issue should be given priority? How would you handle this situation?
Answer:
Several things are going on here:
- In the first reported issue, you didn't have clear indication of impact. If people couldn't trade, that might be more important than producing a report (the 2nd issue), despite what you heard on the phone from either partner. However, it sounded like the report was needed for reconciliations, which some groups depend on for trading. In this case, you need clarification on the issues. Get the business users on the phone again and get find out more. One of the things you have to get really good at in Support is to ensure you really understand the problem. Sometimes, calling the users back and getting clarity is the only way to accomplish this.
- If you're alone in a shift and need help, call and get it. It's better to take a few minutes and escalate, than to try going at it alone. Remaining calm and really thinking about the best approach is a sign of maturity in a Support associate. Wake someone up if you have to. I always tell my guys it's better to wake someone up than to let things fall apart causing financial loss.
- Keep in mind that there are no two issues that are really, exactly the same in terms of urgency. One is usually more urgent than the other. But suppose they were the same and you can't get help. In this case all you can do is work them on a first-come-first-serve basis. You being the sole person on a shift and not having enough bandwidth to handle multiple issues is more a sign of bad coverage (and ineffective management) than anything. Of course, you wouldn't say that in an interview ;)
I hope you find this answer helpful and that it will help you ace your next Production Support job interview!
Wednesday, September 18, 2013
Examples on Increasing Productivity: Part 2 (For Managers)
A reader asked me a good question: "Would you provide concrete examples on how to increase the productivity of my Production Support team?" Although I'll answer with points that would be important for any team, not just Support, I'll provide examples that apply more directly to Production Support groups.
This is Part 2 of this article. Click on the link to go to Part 1.
The first thing you can do as a manager is to set high expectations for yourself and your team. This means two things set expectations and make sure they're challenging enough. One thing that has worked great for me is doing strategic planning at the beginning of every year with all my directs. We keep it simple. We identify things we'd like to improve about the applications we support (The Challenges). Then we come up with Action Items. Action Items define three things (Where we are, Where we want to be, and what we're going to do to get there). Incidentally, there's a technical name for just talking about the challenges and not coming up with what you're going to do to fix them...It's called complaining.
Make sure the action items achievable (yes, you can use SMART goals), but make sure they'll also challenge your team to do their best. Having clear guidelines and defined projects has worked wonders for the amount of work that my groups achieve. People don't come into work wondering what they're going to do. If the BAU work (incidents, service requests, etc.) is low, then it's time to pull out the plan and work on those strategic objectives. I review plan progress on a weekly basis and provide quarterly updates to senior management to ensure my directs' work is being highlighted and they're getting the right visibility level.
Keep a constant eye for ways to maximize the productivity of your team. Just because you have a plan doesn't mean you can't include important items on the go. Also, if something that was previously identified as important, no longer is, then remove it from the plan. Work only on those things that will add value to your group.
As I've said in prior posts, take time to reward and reinforce productive behaviors. Production Support teams go through a lot of stress and team members need to know that their work is not going unnoticed.
Finally, use your metrics to determine areas for improvement and to track how Productive your teams are being. If all the work you've planned to do is not having a positive impact on your Availability metrics, Support effort, Time tracking, etc., then you're not focusing on the right things. Keep the Purpose of Production Support in mind in everything that you do.
This is Part 2 of this article. Click on the link to go to Part 1.
The first thing you can do as a manager is to set high expectations for yourself and your team. This means two things set expectations and make sure they're challenging enough. One thing that has worked great for me is doing strategic planning at the beginning of every year with all my directs. We keep it simple. We identify things we'd like to improve about the applications we support (The Challenges). Then we come up with Action Items. Action Items define three things (Where we are, Where we want to be, and what we're going to do to get there). Incidentally, there's a technical name for just talking about the challenges and not coming up with what you're going to do to fix them...It's called complaining.
Make sure the action items achievable (yes, you can use SMART goals), but make sure they'll also challenge your team to do their best. Having clear guidelines and defined projects has worked wonders for the amount of work that my groups achieve. People don't come into work wondering what they're going to do. If the BAU work (incidents, service requests, etc.) is low, then it's time to pull out the plan and work on those strategic objectives. I review plan progress on a weekly basis and provide quarterly updates to senior management to ensure my directs' work is being highlighted and they're getting the right visibility level.
Keep a constant eye for ways to maximize the productivity of your team. Just because you have a plan doesn't mean you can't include important items on the go. Also, if something that was previously identified as important, no longer is, then remove it from the plan. Work only on those things that will add value to your group.
As I've said in prior posts, take time to reward and reinforce productive behaviors. Production Support teams go through a lot of stress and team members need to know that their work is not going unnoticed.
Finally, use your metrics to determine areas for improvement and to track how Productive your teams are being. If all the work you've planned to do is not having a positive impact on your Availability metrics, Support effort, Time tracking, etc., then you're not focusing on the right things. Keep the Purpose of Production Support in mind in everything that you do.
Examples on Increasing Productivity: Part 1 (For Associates)
A reader asked me a good question: "Would you provide concrete examples on how to increase the productivity of my Production Support team?" Although I'll answer with points that would be important for any team, not just Support, I'll provide examples that apply more directly to Production Support groups.
So, let's start with what you can do as an associate to improve the productivity of your team. There's an implication, here, and that is, that productivity increases are not just the responsibility of managers (though we'll talk about things managers can do). First of all, be open, ask your manager the question "What can we do to be more Productive?" Many times as Support groups, we get bogged down in the day-to-day, tactical, activities and we don't spend enough time thinking strategically. A question like this one, during a team meeting, might spark a conversation with your entire group about the things that can be put in place. Collect those ideas and come up with approaches to get them effected.
Another thing you can do is determine what you can do to increase internal and external client satisfaction. Let me provide an example of each:
Even associates can help when it comes to expense reduction. I was at a company where we used a monitoring tool that cost over $1MM in licensing annually. It was quite feature rich, great graphical interface, etc. But as it turns out, we needed something a bit more basic. A simple dashboards that would display alerts was all that we needed. Most of the Support people in my group had Development backgrounds, so we took on a project to build a monitoring tool. A few months later we delivered the tool and were actually able to replace the vended software. We saved that $1MM in expense.
Increasing productivity might also be defined as stopping low value tasks and doing more productive tasks. I was in a Support group where the monitoring was quite noisy. There were tons of alerts and people would clear them out every day. Day in, day out, clear the alert. Repeat. Doing this is low value. Instead, we cleaned up the monitoring. We put a list together of noisy, false-positives and embarked on a project to clean them up: configuring the tool to ignore some, reclassifying the severity of the alert, removing the alert from the code altogether, etc. The now quiet monitoring tool enabled us to focus on more value added tasks, like building automation and putting together tools to help the support effort.
If you are a manager, click here to go to Part 2 of this article.
So, let's start with what you can do as an associate to improve the productivity of your team. There's an implication, here, and that is, that productivity increases are not just the responsibility of managers (though we'll talk about things managers can do). First of all, be open, ask your manager the question "What can we do to be more Productive?" Many times as Support groups, we get bogged down in the day-to-day, tactical, activities and we don't spend enough time thinking strategically. A question like this one, during a team meeting, might spark a conversation with your entire group about the things that can be put in place. Collect those ideas and come up with approaches to get them effected.
Another thing you can do is determine what you can do to increase internal and external client satisfaction. Let me provide an example of each:
- Internal: I just had a conversation last night with one of my directs. A user had asked a question and it was taking longer than normal to resolve it. The gist of it was that it was a different Support group who should have been handling the query, but somehow it landed on my team's lap. What my team had done was forward the e-mail to the other Support group and there had been no response. My challenge to the team was that we should take more ownership of issues. It would have been better to call the user to clarify the problem. Instead of sending an e-mail, it would have been better to engage the other Support group directly, over the phone, so that a richer conversation could have happened. This would have been a great opportunity to transfer accountability, reassign tickets, convey urgency, etc.
- External: In a prior gig I had, we had many institutional banking customers who connected to our systems to receive prices on financial instruments. If a client was not connected to us, they were also not dealing with us. This means loss of revenue, of course. The went through the logs and found out approximate times that customers normally connected (we didn't have documented SLAs, a problem we inherited). We set up monitoring for each customer and we put a threshold on the monitoring such that, if they didn't connect after a period of time from when they normally did, an alert showed up in our dashboards. This prompted us to call the client and ask them to connect. Many times they didn't know they weren't connected. This small effort increased revenues for the bank and customers really appreciated being notified.
Even associates can help when it comes to expense reduction. I was at a company where we used a monitoring tool that cost over $1MM in licensing annually. It was quite feature rich, great graphical interface, etc. But as it turns out, we needed something a bit more basic. A simple dashboards that would display alerts was all that we needed. Most of the Support people in my group had Development backgrounds, so we took on a project to build a monitoring tool. A few months later we delivered the tool and were actually able to replace the vended software. We saved that $1MM in expense.
Increasing productivity might also be defined as stopping low value tasks and doing more productive tasks. I was in a Support group where the monitoring was quite noisy. There were tons of alerts and people would clear them out every day. Day in, day out, clear the alert. Repeat. Doing this is low value. Instead, we cleaned up the monitoring. We put a list together of noisy, false-positives and embarked on a project to clean them up: configuring the tool to ignore some, reclassifying the severity of the alert, removing the alert from the code altogether, etc. The now quiet monitoring tool enabled us to focus on more value added tasks, like building automation and putting together tools to help the support effort.
If you are a manager, click here to go to Part 2 of this article.
Monday, September 16, 2013
A Lesson Learned
We had reached a critical point in the meeting. For several weeks now, we'd been focused on defining a laundry
list of projects that we were about to embark on. The goal was to standardize Support processes across the
organization to achieve greater efficiencies in terms of: tracking metrics, managing incidents, monitoring and
alerting, etc. You name it, we had it covered. All of the Support processes were to be same across the
organization. Things were going to be much easier for everyone.
Right then, one of the managers declared that he had no interest in doing the work. So, we asked why. Was it a
bandwidth concern? Was it a funding problem? Did she not find value in performing the work? The answer to all
the questions was a No. So what was it we asked. Her response: "My manager simply doesn't care whether I do the
work or not. She hasn't asked me to do it, so I don't think I really need to."
Someone else chimed in and said something quite similar.
There are several things we can learn from this story (true story, by the way). The first is for us Support
managers out there:
Show your teams you care about their work.
Support teams go through a lot of stressful situations. It can be a thankless job. But if you as a manager don't
take the time to acknowledge your group's efforts who will? There is nothing that kills momentum and initiative
more than managers who don't recognize their groups' efforts. For Support teams, not having engaged managers who
recognize the importance of the Support effort can be deadly: People get burnt out. The due diligence in
monitoring goes away. They snooze on alerts instead of reacting aggressively. Or they simply feel too disengaged
to work on those initiatives that can really make things better.
As managers we need to learn to take time to celebrate your teams' accomplishments. Send that thank you note or
two. Gather the troops around and recognize that person who went the extra mile. A small cheer or clap might be
all that's needed to re-energize that team member who used to be great but has fizzled out a bit.
For team members there's something to learn from this story as well:
Do the right thing.
Never stop doing the right thing, just because you don't think your manager cares. There's value in Availability
metrics (this is a report on how well your apps are doing). Be relentless with Problem Management, this is what
makes your applications more stable. Take on those projects that will help make it better all around. Never give
up. Your efforts will be recognized. Who knows, perhaps one day you'll have the leadership of the group
and you can be a different kind of manager.
In the end,We do what we do, not because our manager cares. We do it to enable a business. It takes a special
person to wake up in the morning, know you're going to do Support and still come into the office with a smile on
your face. In many respects the terms Application Support or Production Support don't really do justice to what we
do.
So, keep the goal in mind and keep driving towards it. Your business will thank you for it and
you'll feel much better about those daily achievements that come from Production Support.
Friday, September 13, 2013
The Info You Need When You Need it Most: Runbooks
For this post, I'm going to continue to focus on the knowledge aspect of an application. In particular, I'll talk about Runbooks.
Runbooks should be the first point of reference for anything related to an application. Each and every application you support should have a runbook. Otherwise, it would be like flying an airplane without a manual (for those who didn't catch the reference, every pilot has to use the airplane manual when starting it, no matter how familiar they are with the model).
Runbooks should contain some key information about an application.
The most important section a runbook should contain is a Business Context section which provides the users some idea of the business processes, their criticality and potential financial impact. Most runbooks I've seen don't contain this section, but I like to have this in place. This section should help to further solidify to a group of techies that they don't support some technology or application, but a business instead.
Runbooks should inform the analyst about the Architecture of an application. It should provide an overview of the servers and databases they communicate with. The Architecture section should provide a network context for the application, as well. It should also depict any middleware being used and also provide an idea of other upstream and downstream dependencies.
Another key section for the Runbook is an Administration section. This section should provide the user information about things like how to restart processes, scheduled jobs, breakglass procedures and start/end of day checks.
Likely, the most critical section in a runbook, when it comes to incidents, is a Monitoring and Alerting section. This section of the runbook should provide a list of common alerts and how to resolve them. This section might also contain information about the eyes-on-glass procedures for monitoring the application.
Next in criticality from the Monitoring and Alerting section is the Escalation section. The contact details for Development and Key Business users should be documented there. Also, contact information for key Infrastructure teams and Upstream/Downstream teams should be captured.
A section which provides more detail about how the application works would be an Application Deployment section. This section should contain information like which locations an application is deployed in and what dependencies it has.
The Monitoring and Alerting section should be supplemented with a Troubleshooting section which captures the most common issues, known bugs and limitations.
A Tools section in a runbook which contains the common tools the team utilizes for troubleshooting might be a good thing to document as well. New team members would certainly appreciate having a handy list of the tools their teammates use and perhaps links to downloading/installing these tools should be there as well.
A final word about Runbooks. Do you want to assess your team's proficiency when it comes to application knowledge? Make a bulleted list with each section of your runbook. Pick some topics from each section and make a little quiz. You'll now have a quick and dirty way to find out their proficiency level.
Subscribe to:
Posts (Atom)






