Sunday, May 29, 2011

Jose Bautista's Hot Start

I haven't written a post in awhile, but I am now home for the summer and will hopefully be writing a few posts a week. I wanted to write one today on Jose Bautista. I wrote about him at the end of last season here, and I wanted to do a study on why he is even better than last year.

The major difference between this year's version of Bautista and last year's is his much better batting average. He is still managing to hit a ton of home runs, but after hitting only .260 last season, he is now hitting .353 (all statistics through Saturday's games), good enough for second in the AL. One of the biggest reasons behind this increase in batting average is that his BABIP (batting average on balls in play) has increased from .233 to .321 this year. His career rate of .273 suggests that he was somewhat unlucky last year, and has gotten lucky this year. This may be misleading, as his changed swing naturally leads to more fly balls, usually meaning a lower BABIP. It seems as though last year he was simply trying to hit the ball out of the park, while this year he has become more of a line drive hitter while still hitting home runs. This can be seen in his line drive %, which was only 14.4% last year and is up to 17.1% this year. Line drive % shows how "lucky" a hitter is getting, as line drives are usually end up falling for hits, while ground balls and fly balls are more frequently outs. The increase in LD% shows that Bautista hasn't actually been any luckier this season, he is simply hitting the ball much harder in a higher percentage of at-bats and is being rewarded with a higher BABIP and subsequently batting average.

We can demonstrate what could have happened in previous seasons had Bautista hit as many line drives, leading to a higher BABIP. BABIP is calculated as: (hits - home runs) divided by (at-bats minus strikeouts and home runs plus sac flies). Last year, Bautista had 569 at-bats, 148 hits of which 54 were home runs, 116 strikeouts, and 4 sac flies. If we set his hits total as unknown, we can solve for the amount of hits he would have had with different BABIP. He had a total of 94 non-HR hits last year, and if he had even had his career BABIP of .273, he would have produced 110 non-HR hits. This would have left him with a batting average of (164/569) = .288. If he had this year's .321 BABIP, he would have had an batting average of .322, much closer to this year's .353. The difference in batting average from his career to this season is due to both luck and skill, and unfortunately we can't exactly differentiate the two, but we do know that Bautista has become a better hitter, and is hitting for a higher average at least due to some skill.

So where is the other 30 point difference coming from? The BABIP formula shows us that it comes from the number of home runs and strikeouts a hitter accumulates (also sac flies, but they are minimal and can be ignored). As discussed in the last post, a large determinant of the number of home runs a hitter has is his HR/FB ratio, which is mainly due to luck. Obviously, hitters like Bautista who make a conscious effort to hit fly balls very hard will have higher HR/FB rates, and thus more home runs. The league average is usually around 10.6% (actually, the HR/FB rate is now below 9% for 2011), and last year Bautista ended the year with a 21.7% rate, more than double league average. This year, Bautista has a 31.3% HR/FB rate, which is insane. This means that roughly for every three fly balls he hits, one ends up over the wall for a home run. This is by far the highest rate in the majors this year, with Lance Berkman having the second highest rate at 23.4%, which is 34% lower than Bautista's rate. Even with Bautista's violent swing, this rate is bound to regress at least somewhat towards the mean. This shows that Bautista has gotten somewhat lucky this year, but it is impossible to determine what his final HR/FB rate will be, so we cannot determine exactly how lucky.

The last determinant of BABIP is from strikeouts. Bautista has dramatically increased in this area, decreasing his strikeout rate from 20.4% last year to 17.3% this year. This may not sound like much, but over a full season of 500 or so at-bats, that's 16 more balls put in play, and with a BABIP of .321, five more hits. That's about 10 extra points on his batting average over a full season, simply due to striking out less. What makes this even more impressive is that while he is striking out less, mainly due to the fact that he decreased his swinging strike percentage from 7.7% to 6.7%, he still has the ability to hit the ball extremely hard, actually harder this year. In almost every case, a hitter will sacrifice power in order to make more contact, yet Bautista has managed to become better at both. This is certainly not luck, and we can attribute this part of his increased batting average all to skill.

What does this all mean? I have presented a lot of different statistics, and what I hoped to accomplish was to show that while Bautista has gotten a little lucky in his huge increase in AVG this year, it has mainly been from skill. Although at first you may want to credit luck from his increased BABIP, he hasn't simply had more balls "find holes" this year, he has been hitting more line drives, which show that he has become a better hitter. Yes, his home run total has been somewhat inflated by his incredible HR/FB rate, so maybe we shouldn't expect him to hit 40+ home runs the rest of the season and instead expect him to finish with 50 or so home runs. If his HR/FB rate regresses even all the way to last year (which I don't expect it will), he should still end up with 48 home runs. Finally, he has managed to swing and miss less pitches, which decreases strikeouts and allows him to hit more balls in play. This is due to a systematic adjustment, and is not lucky at all. Although we know Bautista won't end the year hitting .350, with all of these factors we just discussed, we can reasonably expect him to finish the year hitting .320 or so. Many people expected Bautista to slump this year and not be able to hit as well as last year, but he has managed to play even better, and should end up with better numbers this season than last, which is hard to believe.

Wednesday, March 2, 2011

Blue Jays Batting Order

Now that the offseason has winded down and spring training games have begun, it is time to look forward to the 2011 season. An important question for this year (and any year) is what will the batting order look like? This breaks down how exactly a lineup should be constructed to optimize players' talents. Basically, the old-school thoughts on building a lineup (speed at #1, bunter #2, power hitters #3-5, worst hitters #6-9) are mostly incorrect, according to sabermetric research. The article does a nice job of explaining who should hit where and why. It does note that a specific permutation of players in a batting order does not make a huge difference, only about a maximum of one win per season. However, it is a fun exercise to construct a projected lineup.

This post was inspired by this post, which detailed the optimized lineup for the Indians. I want to do the same thing with the Blue Jays. The data I am using is the Cairo projections (found here), which project a player's upcoming season based on weighted average of a player's past few seasons. The statistic that is used to figure out the optimal lineup is wOBA, or weighted on-base average. Two good explanations of the statistic can be found here and here. In short, wOBA combines on-base percentage and slugging percentage into one statistic, scaled to OBP, so it is easy to understand. Why not simply use on-base plus slugging? For one, OPS weighs OBP and SLG equally, while in reality OBP is more important. wOBA is calculated using the actual run values of each event, so it will better predict how much more valuable something is worth rather than simply OPS.

Although it sounds confusing, and the math behind it is, the end result is one simple number which tells you how valuable a player is to his team while hitting. If this sounds similar to Wins Above Replacement, it's because wOBA is used to calculate the hitting aspect of WAR.

Now, onto the data. Here are the splits for all of the Blue Jays involved in the Cairo projections.


Player
Projected wOBA
Vs L
Vs R
Jose Bautista
.373
.367
.375
Travis Snider
.339
.314
.344
J.P. Arencibia
.335
.350
.327
Adam Lind
.331
.293
.345
Edwin Encarnacion
.331
.349
.325
Juan Rivera
.329
.343
.323
Luis Figueroa
.329
.327
.330
Randy Ruiz
.328
.339
.322
Yunel Escobar
.326
.335
.322
Rajai Davis
.325
.339
.318
Aaron Hill
.321
.336
.316
Jason Lane
.312
.323
.306
Chris Aguila
.309
.317
.304
Mike McCoy
.307
.320
.298
John McDonald
.288
.302
.281
Callix Crabbe
.285
.288
.283
Jose Molina
.283
.297
.276
 
I am going to use lineups against both LH and RH pitchers, as the projections include platoon splits. The optimal lineup, according to sabermetrics, is #1, #4, #2, #5, #3, #6, #7, #8, #9 in terms of avoiding outs. So the highest wOBA will be first, then 4th, all the way to 9th.

Lineup vs. LHP

#
Name
Position
wOBA
1
Jose Bautista
RF
.367
2
Edwin Encarnacion
3B
.349
3
Randy Ruiz
1B
.339
4
J.P. Arencibia
C
.350
5
Juan Rivera
LF
.343
6
Rajai Davis
CF
.339
7
Aaron Hill
2B
.336
8
Yunel Escobar
SS
.335
9
Luis Figueroa
DH
.327

Lineup vs. RHP

#
Name
Position
wOBA
1
Jose Bautista
CF
.375
2
Travis Snider
LF
.344
3
J.P. Arencibia
C
.327
4
Adam Lind
1B
.345
5
Luis Figueroa
2B
.330
6
Edwin Encarnacion
3B
.325
7
Juan Rivera
RF
.323
8
Randy Ruiz
DH
.322
9
Yunel Escobar
SS
.322

We can see some interesting things in the two lineups. Seven players appear in both lineups, although maybe not who you would think: Bautista, Encarnacion, Ruiz, Arencibia, Rivera, Escobar, and Figueroa. Rajai Davis and Aaron Hill are only hitting against lefties, and Lind and Snider are only hitting against righties.

The projections are obviously not perfect, if Luis Figueroa is starting every day for the Jays, but they do provide some insight into where certain hitters should hit. With a league average wOBA of .321 last year in the majors, the Jays' lineups should be much better than average again this year, even with the loss of Vernon Wells and John Buck.

Monday, January 31, 2011

30 HRs or 30 saves?

I have done two posts, the true value of a home run and the true value of a save. These posts sprung out of the question: which is worth more, 30 home runs or 30 saves?

We found that the true value of a home run was worth 1.406 runs, and that the true value of a save was 0.11415 WPA. So how do we compare these two variables in different units? Eventually, we want to set a dollar value to each event, but we must first translate each into a win value.

It has been estimated that a win is worth somewhere between 9.5 and 10 runs. There are many different explanations how that was calculated and why it is so, but for simplicity I am just going to accept the argument that 10 runs = 1 win. We can now change the run value of a home run into a win. One home run is worth 0.1406 wins, so 30 home runs would be worth 4.218 wins.

Although it is usually not helpful to sum up WPA, in this case, it is the best we can do to approximate the value of a save. We found that the average save is worth 0.114 wins, so 30 saves would be worth 3.425 wins, using WPA. If we use WPA/LI, the average save was worth 0.0614 wins, so 30 saves would now only be worth 1.841 wins.

We have found out that, mathematically, 30 home runs are clearly worth more than 30 saves. We can now figure out how much each are worth in dollars. It has been estimated each win is worth about $4.5 million on the open market (so each win above replacement will cost approximately $4.5 million to replace, obviously a player with 8 wins above replacement is not going to be paid $36 million per year). So the value of 30 home runs, on the open market, is $18.98 million. This seems to be an unrealistic number, but there are players such as Jayson Werth who hit 27 home runs last year and received a 7-year, $126 million contract (average of $18 million/year) this offseason from the Nationals.

30 saves measured by WPA are worth 3.425 wins, or $15.41 million, and 30 saves measured by WPA/LI are worth $8.28 million. This dollar amount for WPA/LI is much more realistic than the amount for home runs. One example is Bobby Jenks, who compiled 27 saves last year and got a 2-year, $12 million contract this offseason.

So, the answer to the question of 30 home runs or 30 saves has clearly been answered. Home runs are either only slightly more valuable, or much more valuable than saves, depending on your view of relief pitchers. I believe that the math agrees with intuition here, as it seems as though it would be much easier (and cheaper) to acquire a player that will get 30 saves as opposed to a player that will hit 30 home runs. The marginal difference between an average closer (like Frank Fransisco for the Jays) and another pitcher in the bullpen (say, Jason Frasor) is much smaller than the marginal difference between a player like Aaron Hill and a bench player, such as John MacDonald.

In conclusion, I want to show one more example. This is a list of the 18 players who hit at least 30 home runs last year. The average Wins Above Replacement for the players was 4.39. If we look only at Batting Wins (WAR with the defense and running statistics removed), the players still have an average of 3.58 wins. This is a list of the 14 pitchers who saved at least 30 games last year. They have an average WAR of 2.09 wins. I believe that this shows that the pitchers who save 30 games are less valuable to their teams than the players who hit 30 home runs, which we have seen over the past three posts.

Sunday, January 30, 2011

The True Value of a Save

This post will be a little different than the previous post, in that it is much more difficult to quantify the value of a save than the value of a home run. For home runs, we can fairly easily calculate the differences between each one, as there are only 24 different ways for a home run to occur (the base-out states). There are many, many more ways for a save to unfold. There are one and two inning saves; one, two, or three run leads; and a number of different base-out combinations throughout a save attempt.

To determine the true value of a save, we are going to look at all 1,204 saves in 2010, and determine the WPA of each save. A description on WPA can be found here, but put simply, it is the probability of a team winning after an event subtracted by the probability of a team winning before the event. It will show how much a player contributed to his team winning the game. Every save from last year can be found here, sorted by WPA. The true value of a save will then be the average WPA for all saves.

The most valuable save last year was recorded by Andy Sonnanstine of Tampa Bay, with a WPA of 0.662. The least valuable save was recorded by Matt Harrison of Texas, with a WPA of only 0.001 (he actually pitched 3 innings in a blow out game, which is one of the obscure ways a reliever can get a save). The average of all saves last year was a WPA of 0.114. What this means is that the average reliever recording a save will increase his team's expected win probability by about 11% (from 89.6% to 100%).

Unfortunately, there are many debates going on (such as here) as to whether or not WPA is an accurate measurement of a relief pitcher's value. The probability of a team winning when leading going into the 9th inning has not changed whatsoever from 1952 to 2010 (which is pretty amazing!). Naturally, this calls into question the value of the modern day closer. So instead of using WPA, many sabermetricians use WPA/LI, otherwise known as Context Neutral Wins, which is described here. LI is the leverage index of a certain play, as a tie game in the 9th inning will have much more pressure than a play in the 1st inning of a game. Simply using WPA will not account for the context of the situation, so the value of a reliever could be drastically overvalued merely because they pitch in higher-leverage situations.

WPA/LI takes care of this problem by neutralizing the leverage of the situation. As a result, a player's contribution will almost always be less, especially for relievers. If we look at the WPA/LI for all of the saves from 2010, the average is WPA/LI is now only 0.061 (almost half of the WPA value).

So the problem now becomes, which statistic do we use? WPA or WPA/LI? This really depends on your own beliefs. If you believe that closers are really good pitchers who can do things other relief pitchers cannot do, especially in high-pressure situations, then you would want to use WPA. However, personally I believe that closers are only marginally better pitchers than their bullpen counterparts, and as such, are getting some undue credit. So I believe that using WPA/LI is better, especially considering that many closers are failed starting pitchers. However, I will use both statistics in comparing home runs and saves. My next post will finally answer the question of whether 30 home runs or 30 saves are more valuable.

Thursday, January 27, 2011

The True Value of a Home Run

I have been meaning to do a post of the true value of a home run for awhile, but unfortunately I put it on the back burner for awhile until I was asked this question: which is worth more, 30 home runs or 30 saves? In this post, I am going to examine the true value of a home run, and in the next post I will examine exactly how much a save is worth, so I can compare the two.

The data I am going to use is for all teams in the 2010 regular season. The first thing to do is to find the number of home runs hit in each base-out state, which can be found from baseball-reference:
RUNNERS HR_OUTS_0 HR_OUTS_1 HR_OUTS_2
None 1220 811 617
1st 248 318 312
2nd 74 133 147
3rd 9 39 52
1st and 2nd 54 119 142
1st and 3rd 27 49 47
2nd and 3rd 9 30 30
Bases Loaded 23 43 60

The total number of home runs hit last year was 4613, and over half of those were solo home runs. It was very rare for players to hit home runs with no outs and runners on third, as it would usually require a triple, or a double and steal.

The next step is to find the expected runs matrix for 2010 (from Baseball Prospectus):
RUNNERS EXP_R_OUTS_0 EXP_R_OUTS_1 EXP_R_OUTS_2
None 0.49154 0.26151 0.10374
1st 0.85877 0.50512 0.2282
2nd 1.10113 0.67765 0.3215
3rd 1.35798 0.93308 0.34192
1st and 2nd 1.42099 0.88181 0.45503
1st and 3rd 1.80042 1.0982 0.46571
2nd and 3rd 1.96584 1.38849 0.58205
Bases Loaded 2.36061 1.51185 0.77712

We can use these two matrices together to determine the true value of a home run. The equation we will use is: value of a home run = Expected runs at the end of the play - Expected runs at the beginning of the play + the number of runs scored during the play. What this means is that we are taking the expected runs after - before to determine the value of the play (e.g. a leadoff out would be calculated as 0.26151 - 0.49154 = -0.23003, meaning the expected runs for the team in that inning would decrease by 0.23 runs), and then adding the number of runs that were scored.

This matrix shows the true value of a home run for each base-out state. Obviously, when there are no runners on base, the value of a home run will be 1, as the beginning and end states will be the same.
RUNNERS Value_OUTS_0 Value_OUTS_1 Value_OUTS_2
None 1 1 1
1st 1.63277 1.75639 1.87554
2nd 1.39041 1.58386 1.78224
3rd 1.13356 1.32843 1.76182
1st and 2nd 2.07055 2.3797 2.64871
1st and 3rd 1.69112 2.16331 2.63803
2nd and 3rd 1.5257 1.87302 2.52169
Bases Loaded 2.13093 2.74966 3.32662

The most valuable home runs, obviously, are grand slams, as they score 4 runs, while home runs hit with two outs are more valuable than those hit with 0 or 1 out as there will be fewer chances remaining in the inning to drive in the runners or base, thus making the home run more valuable.

Finally, we need to multiply the matrix containing the number of home runs hit by the matrix showing the true value of a home run for each base-out state to find the run values for each base-out state.
RUNNERS Value_OUTS_0 Value_OUTS_1 Value_OUTS_2
None 1220 811 617
1st 404.92696 558.53202 585.16848
2nd 102.89034 210.65338 261.98928
3rd 10.20204 51.80877 91.61464
1st and 2nd 111.8097 283.1843 376.11682
1st and 3rd 45.66024 106.00219 123.98741
2nd and 3rd 13.7313 56.1906 75.6507
Bases Loaded 49.01139 118.23538 199.5972

To find the true value of a home run, we simply add up all of the runs (6485) and divide by the total number of home runs hit (4613) to find the average value of a home run: 1.406 runs. What this means is that the average home run hit in 2010 was worth 1.406 runs for the player's team. We will use this number later to figure out exactly how much each home run is worth in a dollar amount, and whether or not it is worth more than a save.

Thursday, December 23, 2010

Using Markov Chains to Evaluate the Hitting of the 2009 Toronto Blue Jays

This is a paper I wrote for one of my statistics classes last year. The goal of the paper was to figure out the run probabilities for each of the 24 base-out states in baseball (8 base states, 3 out states). I have included the matrices I used to create the run expectancy table if you would like more background info on how exactly the table was created.

Matrices with probabilities:
Markov Chains - Probabilities

Markov Chains Paper:
Markov Chains

Tuesday, December 7, 2010

Predicting MLB Salaries through Offensive Statistics

This is a paper I wrote for my Econometrics class on predicting MLB Salaries through offensive statistics before and after Moneyball was written. It is a pretty long paper (15 pages plus figures, graphs, etc.), but it nicely blends economics and baseball.

MLB Salaries

Saturday, November 27, 2010

Improving a team's Pitching

I have already written two posts on the best way of improving a team, and improving a team's hitting. In this post, I want to do much of the same as the hitting post, but this time on pitching statistics. I am again going to run a linear regression model to determine which statistics are best correlated with pitching performance, which will show us which statistics can be best used to improve pitching.

In this model, instead of trying to estimate runs scored, I am going to use ERA as the dependent variable. Using runs against is a possibility, but since we are estimating the effect of statistics on pitching, and not pitching and defense, using runs against would include the effect of defense, so it is not an appropriate DV in this scenario. We again need to be careful in our selection of independent variables as to avoid collinearity.

Pitching statistics are almost opposite of hitting statistics. Good hitters are generally grouped into two categories: those that can get on base, and those that can hit for power. Good pitchers are those who do not allow very many baserunners and do not allow many home runs. We can measure these qualifications by using the two statistics Walks and hits per innings pitched (WHIP), which measures the average number of baserunners a pitcher allows per inning, and home runs allowed, which will not encompass all extra base hits, but should give us a good feel for pitchers who do and do not allow many home runs that will hopefully be a decent predictor for all extra base hits. Finally, I am also going to include strikeouts as a predictor, because pitchers with high strikeouts rates are valuable, and maybe a pitcher with more strikeouts will allow less runs because he has to rely less on his defense. When we run the regression, we get the following:

                  Estimate        Std. Error     t value    Pr(>|t|)
Intercept    3.2986958     0.4207396    7.840      6.45e-14 ***
WHIP        0.4431320     0.2483891    1.784      0.0753 . 
SO            -0.0014423    0.0001842     -7.829    6.97e-14 ***
HR            0.0117201     0.0008172     14.342   < 2e-16 ***
R2 = 0.551

As you can see from the R2 value, this regression explains a lot less variability than the hitting regression. However, if we replace WHIP by the number of hits and walks given up, we get a lot better regression:

                  Estimate        Std. Error     t value    Pr(>|t|)  
Intercept    -3.410e+00   2.855e-01    -11.942    <2e-16 ***
Hits           3.815e-03      1.547e-04    24.661     <2e-16 ***
BB            2.463e-03      1.414e-04    17.424     <2e-16 ***
SO            -4.711e-05     1.063e-04    -0.443      0.658  
HR            5.298e-03      4.567e-04    11.601     <2e-16 ***
R2 = 0.8906

Now, the R2 value is almost as high as the hitting regressions. All of the variables are significant except for strikeout, so when we take it out of the regression we get the following:

                  Estimate        Std. Error    t value    Pr(>|t|)  
Intercept    -3.5115171   0.1703878   -20.61     <2e-16 ***
Hits           0.0038502     0.0001321    29.14     <2e-16 ***
BB            0.0024598     0.0001410    17.45     <2e-16 ***
HR            0.0053015     0.0004561    11.62     <2e-16 ***
R2 = 0.8905

We can see how insignificant strikeouts were in the regression, because when we remove it the R2 value decreases by only 0.0001 (0.01%). We can now determine which variables impact pitching the most. One more hit given up is associated with a 0.00395 increase in ERA, one more walk given up is associated with a 0.00246 increase in ERA, and one more home run given up is associated with a 0.00530 increase in ERA. Since there are vastly different numbers of hits, walks, and home runs given up, we must also look at the mean of each to determine which will most affect ERA. The mean number of hits given up by a team in a single season is 1469.9, the mean walks is 540.2, and the mean home runs is 172.0. If we multiple the means by the coefficients, we get that, on average, hits will increase team ERA by 5.66, walks will increase ERA by 1.33, and home runs will increase ERA by 0.91. Obviously, we are only looking at statistics that will negatively impact (increase) ERA, so the numbers will look very high, as we are not inputting statistics such as outs or double plays that will positively impact (lower) ERA.

So from the results we can easily see that hits are the statistic that most impacts a team's ERA. So the obviously solution for a team would be to give up less hits, but how? One way would be to acquire pitchers with greater command, possibly leading to those pitchers being able to "nibble" more, making hitters swing at worse pitches. This would probably increase walks, and we already saw that walks also are bad for ERA. A better solution would be to acquire pitchers with a low batting average against and also a low batting average on balls in play (BABIP - although it has been shown that BABIP fluctuates year to year and may not be consistent for any pitchers). Pitchers also want to give up less home runs, but if they can reduce the number of hits against them this should in turn reduce the number of home runs against them.

Thursday, November 25, 2010

Improving a team's Hitting

As a follow up to my last post, which shows how improving a team through hitting or pitching is equally valuable, I wanted to look at how a team should improve their hitting. There are different ways to score runs, and because higher run totals lead to higher win totals, I want to figure out what is the best way to improve a team's hitting performance, thus leading to more runs and accordingly more wins.

Similar to last post, I am going to run a linear regression model, except this time I am going to use "Runs scored" as the dependent variable. Why not simply use "Wins"? If we were to run a regression model with hitting statistics as predictors and Wins as the dependent variable, we will have a much higher standard error, which means that the R2 value will be much lower, showing the the variability in wins is not explained very much by the hitting statistics. So if we have runs as the DV, we need appropriate hitting statistics for the independent variables. This is trickier to figure out then expected, as we cannot have statistics that are correlated with each other, or the regression model will experience "multicollinearity". What this means is that although the overall regression will predict the dependent variable nicely, we will not be able to tell which independent variables are accounting for the variability in the dependent variable. Although this my sound hard to prevent, it can be fairly straightforward, as a quick example will show. If we were to use on-base percentage, slugging percentage, and on-base plus slugging percentage as predictors for runs (or wins), our equation would have a multicollinearity problem. The overall regression may result in a low p-value, showing that we have predicted runs well, but each statistic individually would have a high p-value. We would not be able to tell which statistic is heavily influencing runs as OPS is an extraneous variable, and since OPS is basically measuring what OBP and SLG are already measuring, the best course of action is to remove it from the equation.

This regression demonstrates the collinearity issue:
                   Estimate    Std. Error     t value     Pr(>|t|)  
Intercept     -5.8651      0.2167         -27.070    <2e-16 ***
OBP            22.4591     17.5672       1.278       0.202   
SLG            15.0106     17.5936       0.853       0.394   
OPS            -4.2768      17.5879      -0.243       0.808

R2 = 0.9089

In short, we are going to need to carefully pick our independent variables so they do not experience collinearity. I am going to use the following statistics to try and predict runs: OBP, SLG, and stolen base %. Although there are many different statistics to use, I am using these three because they represent the three basic ways to improve your team's hitting: get on base more, hit for more power, or become a more successful team running the bases. When we run the regression we get the following:

                   Estimate    Std. Error     t value    Pr(>|t|)  
Intercept     -6.0689     0.2246          -27.022   < 2e-16 ***
OBP           17.9231     0.9630         18.612     < 2e-16 ***
SLG           10.7619     0.4947         21.756     < 2e-16 ***
SBperc       0.3956       0.1387         2.852      0.00463 **
R2 = 0.9111


So these three factors explain over 91% of the variability in Runs per Game. Although the R2 value is only slightly higher than the first regression that involved OPS, we can see that all of the statistics are now significant, as opposed to none of the statistics being significant. All three factors have significant effects on runs per game, and OBP has the largest effect. A ten percentage point increase in OBP (e.g. from .350 to .360) is associated with a 0.0179 increase in runs per game. A ten percentage point increase in SLG (e.g. from .450 to .460) is associated with a .0108 increase in runs per game. Finally, a one percentage point increase in SB% (e.g. from 70% to 71%) is associated with a 0.00396 increase in runs per game.

So what does this mean? The best way to increase your team's hitting is to try and score more runs per game. The best way to score more runs per game is to increase OBP. So the best way to increase a team's hitting is to acquire players that will get on base more often, whether it be through a hit, a walk, or a hit-by-pitch. Acquiring players that hit for power will also positively impact a team's hitting, but not as much as players that get on base. So if a team had a limited budget and could only acquire one or two significant players, they should try and acquire those players that can most improve their team's OBP.

Tuesday, November 23, 2010

Improving a Team - Pitching or Hitting?

After the close of the baseball season in early November, teams look to build for next year through trades, free agency, and the draft (which is more for 3-5 years into the future). But how do you build a better team, and more specifically, what will allow you to have a better team? The old question of pitching vs. hitting is always addressed differently by different teams. Last year, the Giants used spectacular pitching with timely hitting to win the World Series, but just the year before the Yankees used a powerful lineup to bulldoze their way to a World Series win. So which is preferable - scoring more runs, or preventing more runs?

In order to answer that question, I looked at every team's statistics in the last 11 years (2000-2010), and gave each team a value for "Playoffs". A 1 meant that the team made the playoffs, a 0 meant the team did not make the playoffs. To estimate hitting, I used the statistic "runs per game", and to estimate pitching I used the statistic "runs against per game" (this really represents overall defense, including both fielding and pitching - to truly isolate pitching a more appropriate statistic would be something like ERA). I then ran a simple linear regression model, with RpG and RApG estimating the binary "Playoffs" statistic. The table below shows the results:

Coefficients:
                  Estimate       Std. Error      t value      Pr(>|t|)   
Intercept    0.28029       0.24027         1.167        0.244   
RpG           0.41595       0.03884         10.709      <2e-16 ***
RApG        -0.41884      0.03694        -11.339      <2e-16 ***
***: significant at the 0.001 level

So the regression line is: Playoffs = 0.28029 + 0.41595*RpG - 0.41884*RApG. The intercept means that disregarding runs for and against, a team will have a 28% chance of making the playoffs. We know that this is close to being correct, as there are 30 teams competing for 8 playoffs spots, so the probability of any team making the playoffs, given that all teams are equal, is 0.2667. The coefficient of runs per game shows that a one-run per game increase in RpG is associated with a 41.6% higher probability of a team making the playoffs. Runs against per game is very similar except it is inverse, as a one-run per game increase in RApG is associated with a 41.9% lower probability of a team making the playoffs. As an example, if a team scores 4.50 runs per game and allows 4.50 runs per game, the probability of the team making the playoffs is 0.2673. If they increase their runs per game to 5.50 (an increase of exactly 1), the probability of the team making the playoffs will increase by 0.416 to 0.6832. If they then increase their runs against per game to 5.50 (again an increase of 1), their probability of making the playoffs will decrease to 0.2644 (a decrease of 0.419).

What this all means is that scoring runs and preventing runs have a very similar impact on a team's success (success defined by a team making the playoffs). Runs against is very slightly more important, but the difference is most likely negligible. Obviously this was a quick study, and only based on a small sample, but we can see that teams should be more concerned with overall talent of acquisitions rather than worry about acquiring only players that will help their hitting or pitching.