Pages

Friday, July 1, 2016

Should We Incorporate Misses into our Evaluation of Goalies?

Back in October DTMAboutHeart posted this interesting article. As usual, please read it if you haven't yet. Basically, the really interesting part of it is his idea of incorporating shots that miss the net into our evaluation of goalies. It's interesting because I think it kind of makes sense (I say this as someone who never played goalie, so take it with what you want). Goalies who position themselves better could force shooters to try to aim for a corner or the side of the net and therefore shoot more pucks wide. And if goalies can force shots to go wide, we should then take notice of it.

But of course we need to look at the numbers. And as you can see if you read it, he shows that there is persistence to a goalies miss%. Of course persistence isn't everything. Something being reliable doesn't imply that it's indicative of "talent". There could be other factors which cause it to be repeatable. Something like team effects, which I covered in my last post. For all we know, these numbers are just measuring the team's ability to cause shots to miss the net and not anything the goalie is doing. So we would need to do an analysis like we did last time. Run the correlation between a goalie's miss% and his team without him. But DTM did this. Look at this tweet


I don't know the exact details behind the numbers. But it looks like it takes a good number of years into account. The sample gets low as you get up there, but the result looks clear. No team effect. Or maybe not?

I say maybe not because I think there is a better way of viewing shots that miss the net. Breaking it up by danger zones. The same Low, Mid, and High made by War on Ice. I guess two thoughts come to mind here. First, War on Ice isn't up anymore. And secondly, they didn't even supply those numbers. And the answer to both of those questions is simple, we can use the play by play files left behind by the War on Ice crew (Of which they deserve a tremendous amount of credit for. It's an amazingly valuable source of info which saves people like me countless number of hours of which we'd spend getting the data ourselves.) to derive these numbers.

Now that we have these numbers we can get to work. Because while miss% as a whole may not show anything, it could be hiding a bias which can only be seen by a more granular look. For example, there may be a real effect in the high danger zone, but because most shots come from the other two zones the signal gets drowned out in regular miss%. Also breaking it up by danger zone is probably a better method in general (as with Sv%).

Before we do that, here's the miss% broken down by Danger Zone and overall (these numbers are from 2007-2008 until 2014-2015 and are 5v5).

Overall_miss%           Low_miss%        Mid_miss%           High_miss%

    27.84%                      30.18%               29.17%                 22.24%

The interesting part here for me is the small difference between Mid and Low. I thought it would be higher. It may be due to the fact that I underestimated how many shots miss the net from the low danger zone by virtue of being blocked and not merely missing. Nevertheless, for those interested, here are the distribution of shots broken down by Zone.

Low_miss%             Mid_miss%           High_miss%
 45.75%                       28.33%                 25.9%


And this is just a little more slanted towards the low danger zone than shot on goals.

Ok, so now we get to the important part of checking for team effects. Before I do so there is one more thing to mention which the astute observer may have picked up. In his same team analysis done above, DTM only uses road shots to control for rink bias (bias due to the scorer's recording of statistics. For example, some arenas inflate SOG.....etc). Some good work on the subject is shown here. Ideally, I would use that method and get the numbers myself, but it's a bit too involved for me. I can work out my own crude version, but I think just road shots will be fine for this analysis. But, of course, ideally we'd be using rink adjusted numbers.

So as I said be before the numbers I'm using are 5v5 and from 2007-2008 to 2014-2015. In order to achieve a better sample for players with more shots, I coupled up the years. Years 2007-2008 and 2008-2009, 2009-2010 and 2010-2011 are paired up and so on. For each pair, I divided it up by team and calculated the goalies who faced road shots for that team. And then, based on what sample I'm using, I run a correlation between what the goalie did for that team (and that team only, any numbers the goalie accumulated for other teams was excluded) and what the team did without him. So I did that and here are the results (the numbers below are r not r^2 and the sample restriction applies for both the goalie and without him):


n
Low_miss%
Mid_miss%
High_miss%
Total_miss%
250+ 
321
.048
.061
.018
.057
400+
256
.03
.08
.032
.034
500+
222
.05
.165
.0275
.068
750+
129
.011
.123
.053
.03

These numbers corroborate what DTM found. There is practically no relationship. The highest relationship we see is with mid danger but it's still rather tenuous at around ~.10. Don't get me wrong, that is something but it is still small. So we can safely say that team effects only play a very small role even when we break it up by danger zone.

Also, we might as well take a quick look at the reliability of miss% by danger zone. Showing no team effect is nice, but we have to know when splitting it up that we aren't just looking at noise. So I did something similar to what DTM did when he split up games into even and odd groups. Instead I did it at the shot level. I lined up all the shots for each goalie and ran a correlation between the missed% for the even group and the odd group. The cutoff was 1000 shots and the numbers here are road shots only (the numbers below are in terms of r).

Low            Mid             High             Total
.221           .333            .466               .303

This is what we generally expect. High, then mid, and lastly low, with total somewhere in the middle. I also ran it for total shots (not just road) and the trend stays relatively the same (with higher correlations of course). Here they are:

Low            Mid             High             Total
.513           .506            .682             .62

 It's interesting because here low seems to be more on par with mid. Maybe rink effects are higher for low? Maybe this is more accurate? I'm not sure. I put more stock in just road numbers but I have to work on getting better estimates. It's also interesting that we get a fair number for low_miss%, as opposed to nothing for Low_Sv%. Nevertheless, both of these show that there is clearly something to these numbers as they are repeatable.


I think all of this (on top of the previous work done by DTMAboutHeart) shows conclusively that goalies have the ability to force shots to miss the net and that team effects are minimal at best. So therefore we should be using it in our evaluation of goalies. And I'd say a better way to look at it is when it's broken down by danger zone. I have some more thoughts on this (so I'd love to hear any thoughts on the subject) and I'll probably post something in the near future on it. Here's a google docs with the numbers-https://drive.google.com/open?id=1yN8fPU4ISVEQpr_GLvr1PTSY7fGNOF1itdNlTY6eXl8

***All Data courtesy of (the now defunct) War on Ice

Sunday, June 5, 2016

Team effects and Sv%

A couple of weeks ago a very interesting post was published by @gentleputsch. Please give it a read. Ok, I'm assuming you read it. What interested me the most, as you can probably tell by the title of my post, was the team effect on Sv%. It was bigger than I expected and kind of caught me off guard. With that said, I had a few thoughts on the post (the main part of his post on considering shorthanded Sv% is another thing in itself so I'll leave it out for now) and I want to rerun the numbers a little differently.

First off, running the data just on the past three years seems odd. For all we know, running it on the previous three year stretch could bring back nothing. So we should really incorporate more data. Also, at first I didn't understand why he chose a three year period, but after thinking about it I agree. If we look at one year, we'll end up having a lot goalies who play 50 to 65 games, not leaving much of a team sample. This is lessened for a two year period and even better for a three year (I guess one could actually do it for 2 years but 3 seems like a better choice). With all that said, let's see the numbers.

Just in case some people don't remember, these numbers here are 5v5. The data here is from 2007-2008 until this year. I divided it up into three, three year periods: 2007-2010, 2010-2013, 2013-2016. To be included, a goalie needs to have played all three years with the same team. But instead of just choosing the goalie with the most minutes for each team, I chose goalies with at least 1250 total shots faced. It's arbitrary, I know. But I wanted to get at least ~50 games for each goalie. The total amount of goalies is 84 (so choosing this over the most minutes doesn't make a real difference, we end up with basically the same players anyway). We then run a correlation between a the player's save percentages and that of the team when he's not on the ice. The following numbers are r (not r^2).

n        Low_Sv%    Mid_Sv%     High_Sv%

84          .25              .01         .264


The numbers here seem a bit different than what @gentlepush published. Low_Sv% is about the same, but High_Sv% is lower and Mid_Sv% is nonexistent. It's kind of odd that we get something for low but not mid. Running the numbers for 1 year and for 2 year spans bring back just about the same numbers for mid and high Sv%, Low was practically zero for those years but for some reason gets a spike here. Either way Low_Sv% doesn't matter because we already know that Low_Sv% means almost nothing. As I estimated in my last post it needs about ~13000 shots to regress it (and it's actually slightly higher because of the team effect). That means when a goalie has logged about that many low danger shots we know about half his talent. So we shouldn't even bother with it. With that and Mid_Sv%, which is in the clear, out of the way we now have to deal with High_Sv%.

So the question here is: How do we account for the team effects? I think the first step is thinking about what makes up a player's High_Sv% (or to be fair Sv% in any zone). We know a player's own skill matters, some luck, and as we covered here team effects. So it's kind of like this:

Observed High_Sv%= Talent + Luck + Team Effects

With this in hand, we can estimate the spread of High_Sv% due to talent.

Observed var =  Talent var + Luck var + Team var

The var here is variance. Basically the spread in observed High_Sv% is made up of the spread in talent, luck, and team effects. So let's get some numbers down (those familiar with Tom Tango's work will recognize this type of equation):

One Standard Deviation of observed High_Sv%= .0178 (I thinks it's important to remind everyone I'm doing these numbers over a three year period, for 1 or 2 year the observed SD would be higher).

Luck= p*(1-p)/ Avg # of Shots (this would be the binomial variance)

p here is the average High_Sv% over the period, which is .832. So.....

(.832)(.168)/798.5 = .0000175 (this is the variance, one SD is therefore .0132)

Lastly, team here would just be our correlation multiplied by the observed spread. So we get one SD of Team effect is .0047.

So let's plug it in.

.0178^2 = Talent^2 + .0132^2 + .0047^2

Solving for Talent we get...... SD of Talent = .011

Now all these numbers are nice, but I don't think anyone really cares. But why does this all matter? Because our estimation of a goalie's High_Sv% talent was previously overinflated. We incorporated the team effects into the numbers. And now that we, presumably, got rid of it we could get a better estimate (which I tried two weeks ago until I saw....) as to how much do we need to regress High_Sv%. Mid_Sv% is the same, Low_Sv% doesn't matter, but High_Sv% needs to be adjusted. As usual Tom Tango provides a solution (it seems like his new and old blog are filled with dozens of ways to regress, each a little different)

So let's do it:

Talent SD^2/Observed SD^2= .378 (this would be r^2)

And now.............. (1-.378)/.378 *799 = 1314 Shots

My last estimate of how much to regress High_Sv% was 1121 shots. Now it's about 200 more. Makes sense to me.

At the end of the day, it seems like we were mistaking a little noise for signal in a goalie's High_Sv%. The SD of Talent is really smaller and we just have to regress a little more. I wonder what the numbers would look like for a goalie metric which instead of dividing up by danger zone assigns a probability to each shot (there are a couple of these models out there). I'd assume the team effect is smaller there.

***All Data courtesy (of the now defunct) War_On_Ice






Tuesday, May 31, 2016

Regressing Sv% by Danger Zone

Goalies aren't easy to evaluate. Or as some might say, they are voodoo. Introduced some time back by the war on ice crew(which will soon be defunct), and used by Nick Mercadante in his Mercad60, they split up shots into low, medium, and high danger based on the position of the shot and whether it was a rebound or a rush shot. And it's been used quite a bit. I've been thinking about it lately, and strangely, I've never actually seen anyone regress it. So I did so and I thought I might as well share it.

What we would expect, from intuition and what we know, is that: Low-Danger is almost all noise, Mid-Danger has a signal but it's weak, and High-Danger is the best and has a fair signal. Of course, we need to confirm this. My first to idea to test all this was to do it the conventional way. Take all goalies and line up all their shots. And run odd-even correlations until r=.5 (that being the correlation coefficient). The amount of shots needed to reach r=.5 would be how much we regress (and we would do so at league average). But the problem is that Sv% (in all zones) take time until a good signal is reached, and there are only so many goalies who have faced a lot of shots. We start to run out of goalies as the sample is too small (I was actually able to get a read on High-Danger Sv% and found it reached r=.5 at 1050......this makes sense as we'll see soon).

Thankfully, there are other ways. On way, which is easier and more convenient, is shown here by Tom Tango. I strongly recommend you read that before continuing (and I also recommend reading his blog in general). Ok (I'm assuming you read it), I ran the numbers on the last 4 years on goalies who faced at least 750 total shots. Here's what we find (It goes without saying....this is all Even Strength numbers). Also I only included sv% for the hell of it and for a self-check (I know it should be about 3000). The best way is regressing each zone independently.

                                  n             Sv%          Low-Sv%        Mid-Sv%           High-Sv%

SD of z Scores:           76            1.307         1.136                1.033               1.298

# of shots regressed:   76            3534          3915                 10127               1001


First thoughts from these numbers is that high danger and regular Sv% (it's near Tango's estimate) make sense. But low and mid seem to be in the wrong spots. Flip it and it makes sense, but it doesn't match up. So what I did was run it on the 4 years previous to that (2008-2012). And what we find is:

                                  n               Sv%          Low-Sv%        Mid-Sv%           High-Sv%

SD of z Scores:           71            1.457           .942             1.268             1.264

# of shots regressed:   71            2638          -11565           1352               1400

For this one Sv% makes a little more sense and High-Danger stays about the same. But as you can easily tell Low-Danger is fucked up and Mid-Danger seems too low. So what we'll do is run them together. We'll just combine the data and see what we find (One note: The percentages used now is for the entire eight year period not each four individually).

                                  n              Sv%          Low-Sv%        Mid-Sv%           High-Sv%

SD of z Scores:           147             1.4            1.046               1.16              1.295

# of shots regressed:   147             2838          12933            2165               1121

And these estimates make the most sense. Sv% is near the 3000 mark, Low-Danger is pretty much just noise, High-Danger is the best and has a fair signal. And Mid-Danger is a weaker signal.

Lastly, I'd like you to think of the spread for each danger zone sv%. Over the past three years the observed Standard Deviation for each one is: Low- .69%,  Mid- 1.317%, High- 2.232%. And what we know is that the amount of noise decreases as we move to the right. So not only does the observed spread increase as we move to High-Danger but we can attribute more of that to talent (the standard deviation due to talent is higher). I think this is important in showing why we should be mainly focusing on High Danger sv% when it comes to evaluating goalies (obviously not only focus on it, just mostly) as it has the biggest spread in talent.

***All Data courtesy of War On Ice

For this who care.....the league average sv% for each zone (and in general) over the past 4 years:

Sv%    Low-Sv%    Mid-Sv%     High Sv%
.923     .9738           .927          .8347





Friday, April 15, 2016

To Secondary Assist or not to Secondary Assist

Just over five years ago Eric Tulsky, who now works for the Carolina Hurricanes, published this article on assists. As you can see, he advocates dropping the secondary assist due to the amount of randomness in it (to be fair, I don't know his stance after publishing that). This is an older article, so I feel like I must as well update the numbers and the analysis and see what we find.

So I'm going to be using numbers from the past six years (2010-2016). I'm just going to do a simple year over year analysis like Eric did, but I'll split up between forwards and defensemen (I tried splitting up forwards between center and wingers but there was virtually no difference). Also the minimum amount players needed to have qualified for the analysis was 400 minutes for both years.

Repeatability

n
Position
A160
A260
A60
1402
Forwards
0.36
0.19
0.43
836
Defensemen
0.20
0.03
0.24


Here are the correlation coefficients for each metric. And I would say this all makes sense. First assists for both are more reliable than secondary assists. More so for forwards than for defensemen, for whom it's almost virtually random. And this is reflected in the total assists numbers. For defensemen it's only slightly higher than first assists. For forwards we see a a little bit more there. 

Of course this isn't it. Let's split up between players who stayed on the same team versus those who switched teams. As Eric T. noted, teammates play a role in assists, so numbers on players who change teams might be closer to the truth. So let's see:


n
Repeatability
A160
A260
A60
1020
Forwards-same
0.36
0.21
0.44
382
Forwards-Diff.
0.32
0.10
0.35
Difference
-0.04
-0.11
-0.09
604
Defensemen-same
0.24
0.12
0.26
232
Defensemen-Diff.
0.06
0.02
0.11
Difference
-0.18
-0.10
-0.15

And you can see that in all categories, players who stayed on the same team showed better persistence. Let's look at forwards: For forwards, first assists goes slightly down. But you really see it in secondary assists, where it takes a bit of a tumble. And this of course results in assists as a whole grading out worse for forwards who change teams. Defensemen, on the other hand, really have their numbers lose a lot. Secondary assists are virtually random when changing teams. And surprisingly first assists take a big hit too. I really didn't expect that much of a difference. 


One final thing though, let's look at what better predicts next years A60. First assists or total assists.

n
Predictivity
A160
A60
1020
Forwards-same
0.40
0.44
382
Forwards-Diff.
0.33
0.35
1402
All Forwards
0.39
0.43
604
Defensemen-same
0.27
0.26
232
Defensemen-Diff.
0.08
0.11
836
All Defensemen
0.23
0.24


As you would imagine total assists edges out first assists in all cases except, oddly, defensemen who stay on the same team. I won't put much thought into that, it's nothing. Also, the edge in each case is really small. I would imagine that this being only one year worth of data contributes to this. As you would imagine, after getting a few years of playing time we can make a better estimate of secondary assists.

Conclusion

The numbers here are, overall, close to the one's shown by Eric T. five years back. And I think the best bet when looking at assists is focusing primarily on primary assists. But secondary assists still matter a little. One year of assists tells us more than primary assists. Not by much, but there is something. And given a few years, we'll get a better judge of a players secondary assist "talent". Just looking at the leaderboards for secondary assist over the past few years will tell you that it means something. Secondary assists may contain a lot of randomness, but they still matter. The gain over primary assists is minimal, but they shouldn't just be discarded as noise (just mostly noise).

**All data courtesy of Corsica.hockey


***Update:

This is a good article on secondary assist-http://fivethirtyeight.com/features/some-nhl-stars-get-more-assists-at-home-than-they-deserve/. The next step would probably be adjusting for rink bias. Also this
 is a good chart showing how small the spread of talent is in secondary assists. Even though it would have been better to do a separate chart for forwards and defensemen.





Wednesday, April 13, 2016

Estimating Shooting Talent

  I guess the first question here is, what do I mean when I say shooting talent? That's a fair question. To start off, let's think about shot quality. Shot quality is the probability of a shot going in. This quality is made up of a lot of different things. The distance of the shot from the goal, the angle, the type of shot, if it's a rebound, if it's a rush shot, the score.....etc. Another, sometimes overlooked, component is the the actual shooting talent of the shooter in question. Meaning in hockey parlance, "How good is his shot?". But how do we do this? (Dtmaboutheart's model tried to take this into account in his expected goals model here by regressing sh%, but I disagree with that method)

I think the best way to start to answer this, is to think about what makes up a player's shooting percentage. This as opposed to looking at sh%, because straight sh% isn't exactly shooting talent. A player could have a high sh% because of the quality of shots he takes or because of his actual shooting ability. Nevertheless, I think it's roughly like this:

Observed Sh%= Shot Quality* + Shooter Talent + Randomness

*When shot quality equals to the stuff I talked about in the first paragraph

So what we need to do: Is estimate the shot quality for each player, and regress for randomness and what should be left is our best estimate of a player's shooting talent. And the way to do this, is I think simple. First we'll estimate the quality using the expected goal model developed by Emmanuel Perry over at Corsica.hockey. And then we just do is:

Goals/Expected Goals= Shot Multiplier

Let's think about this. Let's compare two players: Both have the same expected goals but one exceeds in actual goals. So why is one player doing better than the other. Is this because of shooting talent or just randomness? If it's something real we would expect this measure to show some repeatability. If a player consistently scores more goals then we would expect, something is probably going one. But how much? This is obviously a sensitive measurement, so we need to be careful if we are measuring anything real or just pure randomness. So we'll run a regression.

 First we have to split up forwards and defensemen. Those are two different positions and we should obviously expect forwards to have more talent here. And I'm going to run three regressions for each group. The first is a regular year over year regression, limited to players who have 500+ TOI in both seasons. But one year for this isn't the best judge. There is bound to be a good deal of randomness here, and it'll probably take a few years to get a better read. So my other two regressions will be 2 years vs. 2 years, and 3 vs. 3. The cutoff being 1000, and 1500 minutes respectively. These numbers are arbitrary and one can play around a little, but it still leaves us with a fine sample size for the analysis. Also there are obviously better ways to conduct this analysis, but this is a quick and easy way to do so. So here you go (the numbers are 5v5 data from 2007-2008- 2015-2016).



n
TOI
Forwards (r)**
Defensemen (r)**
1974/1213
500+ 
.17
.04
1333/784
1000+ 
.30
.10
782/452
1500+
.39
.13

** r being the correlation coefficient

So let's see what we got here. First off, on the left in the "n" column is the sample size ordered by forwards then defensemen for that particular regression. What we see is simple and intuitive. There is a good deal of randomness in the data and we see more of a signal for forwards than we do for defensemen. For defensemen, there doesn't seem to be much there. Even after three years it still regresses 87% to the mean. It doesn't mean it doesn't matter at all....but it doesn't seem to make much of a difference (specifically when you take into account the amount of goals you would expect these guys to score anyways).

Forwards are obviously a different story. As you can see, we do get a signal through. There still is a fine amount of noise (especially for one year). But considering, it's not bad. We have to be careful with our data and regress it. Given a few years, we can start to make an estimate of a player's shot multiplier.

That's not it though. There's still one question: What do we regress to? As shown first in this article by Eric T., we can use coaches decisions to find the right mean to regress too. Because players who average more TOI per game tend to shoot higher. Make sense. Coaches know how to evaluate talent, and will therefore give more ice time to players who are better skilled. But we can't just plot TOI/G vs. Sh% (as he notes). I would say most of the players given more ice time will be the more talented one's, but there's also the problem of coaches riding the hot hand. For example, imagine a player with regular shooting talent. But he goes off (like many do) for a portion of one year and starts scoring a ton of goals. It's possible that the coach will then bump him up a line and play him more. So getting lucky and over-exceeding one's talent will therefore be reflected in TOI/G. Because some of those players with high TOI/G will be guys who experienced a decent portion of luck and are now getting "rewarded" with more ice time. So it'll be more extreme then it actually is.

What we should do is plot TOI/G in year 1 vs. Shot Multiplier in year 2. This is because year 2 performance is a better estimation of the players talent. In year two those players who got "lucky" and were given more ice time will on average regress towards the mean, so the bias won't be there anymore (remember this also works the other way with players who got unlucky and were given less ice time). So let's look at it:





This is just for forwards. I'll spare posting the graph for defensemen as it's just a straight line. That's in line with what we know so far. Now this graph for forwards shows a clear upward slope. Those players who who get about 13+ minutes would be expected to regress to higher mean, and those below 13 minutes to a lower one. For any TOI/G all you have to do is stick it in the equation to figure out what we would expect.

Ok so there are a few stuff here, so let's look at a working example. Let's compare two players at the extremes: Steven Stamkos and Tanner Glass over the past three years.

                             G60          ixG60          Shot Multiplier
Steven Stamkos:     1.23          .86               1.46

Tanner Glass:         .23            .40              .60

So Stamkos scores about a goal more per 60, but get knocked down about ~.4 goals per 60 in expected measures. So his multiplier is pretty big. Glass, on the other hand, is projected to score a little under double the amount of goals than he actually did. Resulting in a small multiplier. Let's now regress them.

                               TOI-Mean         Regressed Multiplier      
Steven Stamkos:         1.05                    1.21

Tanner Glass:             .86                       .75

So the second column is the mean multiplier we would expect based on the players TOI and should therefore regress to. And the second column regresses their multiplier towards the mean using the numbers discussed in the first chart (it's three years so they are regressed 61% to the mean). Let's look at the final result

                              Regressed Multiplier x G60
Steven Stamkos:         1.04

Tanner Glass:              .30


So the last step is just multiply expected goals by the multiplier. Remember, the multiplier is how much we would expect a player to exceed or fall short of their expected goals. So we would expect Stamkos's G60 to be a higher by a factor of 1.21, and for Glass's to be 75% of his expected G60. I'd like to add that this is only on three year data. If we used players career data and regressed that we would get a better estimate of their shot multiplier. Either way Stamkos here goes up .18 goals per 60 and Glass down .1. These are meaningful differences. These two players are at the extremes so we'll see some of the biggest differences with them, but it still should be looked at.

Conclusion

As expected goal measures become more popular, it's important to remember to "shot talent". Sh% includes both location measures and actual shooting talent. And I would say the best way to measure that in the present time is: Comparing it to results and regressing and using a coach's judgement. It's important to be careful with this, and while it won't capture everything and for a lot of players won't mean much, it's a better way of getting the full picture.

***All Data Courtesy of Corsica.Hockey