Showing posts with label 8th Grade Statistics. Show all posts
Showing posts with label 8th Grade Statistics. Show all posts

Thursday, March 17, 2011

Sabermetric Bracketology 2011

I've been applying sabermetric (or, more appropriately, APBRmetric) principles to my NCAA bracket selections for the past three years. In the contest that uses upset points, I was the big winner the first two years, then fell to a disappointing fourth place last year. So, I've been looking to change my methodology for this year.

Fortunately, master prognosticator Nate Silver came up with his own bracket forecast this year, which I was able to modify for my upset contest's scoring system.

To recap the upset scoring system, a team gets a point per round for winning (one point for winning in the round of 64, two for the round of 32, up to six for winning the championship game), plus bonus points for defeating a lower ranked seed. Bonus points are the difference between the two seeds. For example, if a 12 beats a 5 in the first round, they get one point for the win, plus 12 - 5 = 7 bonus points, for a total of 8.

I then took the potential points each team gets for the win and multiplied by the probability of winning to get the team's expected value for each game. In the first round, this was easy - it's just the probability of winning times the difference in the seeds. For example, Notre Dame has a 91% chance of beating Akron. So ND's expected value is .91*1 = .91, while Akron's is .09*(1+15-2) = 1.26.

In later rounds it's a bit more complicated, since the number of points a team gets for winning depends on who they play. If Marquette wins their first round game, they might play Syracuse, a lower seed (which would earn them bonus points with a second round win), or Indiana State, a higher seed. So, the calculation adds a layer.

In the Marquette example, they have a 90% chance of playing Syracuse in the second round (since that's the Orange's probability of winning the first round game) and a 10% chance of facing the Sycamores. Marquette has a 16% chance of advancing past the second round, and would earn 2 points for beating ISU and 10 for beating Syracuse. So the calculation becomes .16 * (.90*10 + .10 * 2) = 1.47.

I propagated this formula for every team for every round, until I got a total expected value for all six rounds for every team in the tournament. From there, I started pairing off the teams into games. I determined the winner by choosing the team who had the greater expected value over the remainder of the tournament, figuring that that would help maximize my value. I'm not sure if this is the best method - some sort of Monte Carlo simulation would probably be the ideal - but it was at least something I could take for a test run.

Some interesting observations along the way:
  • Silver's bracket has Notre Dame as the "worst" two seed, with just a 1.8% chance of winning it all. Of course, that may be more because it likes Purdue so much, and not because it dislikes Notre Dame.

  • Upset points considered, Clemson is a heavy favorite over West Virginia. Of course, before I could write Clemson into the second round on my bracket, I had to make sure they first won their play-in game.

  • As with my past methods, the odds still favor putting all the #1 seeds in the Final Four. Since I'll be playing against a bunch of Ohio State homers in my upset contest, I'm hedging my bets by having Duke beat them in the semis and beating Kansas in the finals.

  • I also used Silver's bracket to fill out a bracket that had a more straightforward scoring system. But I also took some liberties, putting Notre Dame in the finals, where they'll lose to SDSU. Apparently I'm not as big of a homer as Luke Harangody, who has the Irish winning it all in the "celebrity" bracket he did for Fox Sports Ohio.

Wednesday, January 26, 2011

2011 Cleveland Indians Lineup by The Book

Last year, Manny Acta made a splash by dropping Grady Sizemore to second in the batting order. This year, he's considering moving him back to leadoff. Is either the right move? And how should the rest of the lineup look?

The Book, one of the best sabermetric books you can find, did extensive work on lineup construction. Their main conclusion was that lineup order didn't matter too much, but it can be optimized for marginal gains. The Book's findings are summarized very well in this Beyond the Boxscore post.

To get the stats for Cleveland's upcoming season, I used the Cairo Projections, which are described (and available for download) here. The nice thing about version 0.5 of this years Cairos is that they include lefty/righty splits. It uses wOBA, which is decribed in detail in the new Frangraphs library. As you can see, wOBA is scaled to be comparable to batting average, with a .321 wOBA being the league average in 2010.

First, here's how the Indians lineup should look against lefties. I took the top nine players in terms of wOBA against lefties, and fortunately things worked out nicely in the field.
ordernameposwOBA
1Shin-Soo ChooRF.343
2Matt LaPorta1B.351
3Shelley DuncanLF.332
4Carlos SantanaC.346
5Austin KearnsCF.342
6Jayson Nix3B.327
7Asdrubal CabreraSS.326
8Travis HafnerDH.326
9Jason Donald2B.325


The glaring omission, of course, is Grady Sizemore. Cairo projects Sizemore to have a wOBA of only .309 against lefties. But if you insist on playing him (both in the name of fan interest, and so Kearns doesn't have to play center), you can remove Hafner from the lineup, DH Duncan, and move Donald up to eighth with Grady batting ninth.

Some other items of note:
  • Everyone in this lineup is projected to hit above a .321 wOBA. That's nice, but .321 was the average in 2010 against all pitchers. The average against lefties in 2011 may be higher or lower.
  • Indians fans should be especially pleased to see such a nice number for Matt LaPorta, especially after his struggles at the plate these past few years.
  • LaPorta and Santana have very similar numbers, but Santana has a slight edge in power, giving him the fourth spot over LaPorta. While Choo also has very good power, his on base percentage is just too good to put anywhere but first.


Now, the lineup against righthanders. Unfortunately I wasn't able to take just the best nine hitters this time. Michael Brantley and Travis Buck both rated ahead of Jack Hannahan. Brantley, Buck, and Duncan all rated ahead of Nix and Donald as well. But somebody has to play second and third base.
ordernameposwOBA
1Shin-Soo ChooRF.390
2Carlos SantanaC.359
3Matt LaPorta1B.332
4Grady SizemoreCF.363
5Travis HafnerDH.342
6Austin KearnsLF.322
7Asdrubal CabreraSS.318
8Jack Hannahan3B.309
9Jayson Nix2B.307


If you don't think Jack Hannahan is going to break camp with the Tribe, feel free to move Nix up a spot in the order and plug Jason Donald's .303 wOBA into the nine hole.

Notes on this lineup:
  • Choo blew everyone away in both on base percentage and slugging. But I chose to hit him leadoff, just to give our best hitter as many at bats as possible.
  • Believe it or not, Sizemore is expected to have better slugging numbers than Santana, and Santana better on base numbers than Sizemore. That's why Grady is hitting fourth and Carlos second.
  • Cabrera, Nix, and Hannahan/Donald will need to be good with the glove to make up for their below-average projections. Other than that, though, this isn't too bad a lineup.


Finally, for those interested, here are the numbers for a few key players who failed to crack either lineup:
namewOBAvs Lvs R
Michael Brantley.310.291.316
Travis Buck.306.288.312
Luis Valbuena.300.286.302
Trevor Crowe.289.283.290
Adam Everett.268.282.264

Wednesday, August 25, 2010

2010 Fans Scouting Report

Once again, TangoTiger is crowdsourcing his defensive scouting reports. If you follow baseball closely enough to have an educated (or semi-educated) opinion on players' defensive talents, be sure to stop by and enter your thoughts.

Wednesday, June 16, 2010

Cleveland Indians Monthly Statistical Analysis

A (hopefully) monthly look at the Tribe's performance from a variety of statistical angles. Suggestions for additional stats and categories are welcome - just let me know.

Overall

Pythagorean Record

One of the most popular judges of a team's performance is their Pythagorean record. "Pythag" is so named because the formula used looks similar to the Pythagorean formula used on right triangles.

Pythag uses runs scored and runs allowed to determine a predicted record. Generally, teams that overperform their Pythags are considered lucky and those that underperformed are considered unlucky. However, small ball teams like the Angels and Twins consistently overperform their Pythagorean records, for reasons saberists have trouble explaining.

So far, the Indians Pythagorean record (which can be found on MLB.com's standings page by turning on the "X W-L" category) is exactly the same as their actual record, 21-34.

Recently, Fangraphs suggested that using runs scored/runs against in-season isn't the best for Pythag. Instead, Base Runs should be used. This blog quickly picked up that idea and ran with it. However, even using base runs, the Indians can't shake their 21-34 record. Projecting that out over the course of the season, the Tribe is looking at a 62-100 record.

Playoff Odds

Several sites offer odds of a team making the playoffs. One such site is CoolStandings.com, which currently has the Indians at a 0.5% chance to win the division and less than a 0.1% chance to win the wildcard.

Baseball Prospectus isn't as optimisitic. Their projections, explained at the bottom of the linked page, have the Indians at 0.14% to win the division and 0.006% to win the wildcard. However, their projected standings do have the Clevelanders finishing at 70-92.

"Wins in the Bank" and the Gambler's Fallacy


Several sites also post preseason standings projections. This is where the "gambler's fallacy" comes in. BaseballProjection.com has the Indians going 81-81 on the season, yet they have gone 21-34 so far.

Someone following the gambler's fallacy would think that the Tribe will go 60-47 the rest of the season to finish 81-81. In fact, Chone's projection should be thought of as a winning percentage, not a definite number of wins and losses.

That is, we should accept the fact that the Indians are 21-34, and assume they will play .500 ball (81 wins/162 games) the rest of the way. That means the Indians should expect to win 53-54 of their remaining 107 games, for a final record of (at best) 75-87.

Pitching

BABIP

One of the most consistent statistics in baseball is a pitcher's batting average on balls in play (BABIP). The more a pitcher throws, the more his BABIP will settle in the .290-.300 range. (Thank you to the commenter who corrected me last month. This is true for nearly all pitchers, with only rare exceptions like Jim Palmer, who had a great defense in Baltimore.

That being said, if you see a pitcher with a high BABIP early in the season, you can expect that BABIP - and his overall performance - to improve. Conversely, pitchers with a low BABIP can expect an increase.

According to Fangraphs, Jake Westbrook and Chris Perez are already in the .290-.300 range. Those two have put up decent numbers and should expect to continue at the same pace. That may be good news for the Indians front office, who could potentially swap Westbrook for some prospects within the next two months.

Cleveland's two best starters, Fausto Carmona and Mitch Talbot, have been beneficiaries of below-average BABIPs and can look forward to some minor regression. The same goes for relievers Tony Sipp, Joe Smith, and Frank Hermann. In fact, Sipp may have already seen the start of that regression during Cleveland's most recent road trip. Smith and his 7.71 ERA, meanwhile, won't take kindly to the notion that he's been lucky in a good sense. Hermann has been lights-out in a limited Major League debut, but he can't keep up this level of talent forever.

Meanwhile, a majority of the Tribe's staff should see their luck increase. David Huff and Aaron Laffey are due for some minor improvement, while the embattled Kerry Wood, Justin Masterson, Rafael Perez, and Hector Ambriz should see drastic drops in their BABIP as the season progresses.


Fielding Independent Pitching

While BABIP looks only at balls put into play, fielding independent pitching (FIP) does the opposite. FIP looks at walks, strikeouts, and home runs - the three things that are supposedly the pitcher's responsibility. In other words, FIP looks at those things that can't be affected by the quality of the defense.

FIP is calculated on the same scale as ERA, so the two can be compared easily. If a pitcher's FIP is lower than his ERA, he can expect improvement, and vice-versa. Of course the major caveat here is that if a pitcher plays with the same defense all year, how much can it really improve?

Half of the Indians staff has a FIP within 1.00 of their ERA. Chris Perez, Frank Hermann, and Mitch Talbot are greatly outperforming their FIP, while Justin Masterson, Aaron Laffey, Rafael Perez, and Kerry Wood are greatly underperforming it.


Hitting

Lineup Analysis

BABIP can be used to study a hitter's luck. But unlike pitchers BABIP, hitter's BABIP normalizes to the hitter's past performance, not an overall league average. I'll leave that as an exercise to the reader.

Instead, for hitters we'll again refer to Dave Pinto's lineup analysis tool. Plugging in the top nine Indians in terms of OPS, here are the results. How plausible is the lineup? Not very, unless you can convince one of the outfield/first base/DH-types to catch.

This theoretical lineup would score 5.138 runs per game, much better than the 4.036 the team is currently scoring. That's an extra 178.5 runs for the year, or 18 wins. Of course, that's based on some small sample sizes, and a less than stellar defense.

Wednesday, May 12, 2010

Cleveland Indians Monthly Statistical Analysis

A (hopefully) monthly look at the Tribe's performance from a variety of statistical angles. Suggestions for additional stats and categories are welcome - just let me know.

Overall

Pythagorean Record

One of the most popular judges of a team's performance is their Pythagorean record. "Pythag" is so named because the formula used looks similar to the Pythagorean formula used on right triangles.

Pythag uses runs scored and runs allowed to determine a predicted record. Generally, teams that overperform their Pythags are considered lucky and those that underperformed are considered unlucky. However, small ball teams like the Angels and Twins consistently overperform their Pythagorean records, for reasons saberists have trouble explaining.

Likewise, Eric Wedge's teams traditionally underperformed their Pythagorean record. That may be because those teams sprinkled 15-run wins between seven-game losing streaks. So, it will be interesting to see how Manny Acta's teams perform against Pythag.

So far, the Indians Pythagorean record (which can be found on MLB.com's standings page by turning on the "X W-L" category) is exactly the same as their actual record, 11-18.

Playoff Odds

Several sites offer odds of a team making the playoffs. One such site is CoolStandings.com, which currently has the Indians at a 1.4% chance to win the division and a 0.5% chance to win the wildcard.

Baseball Prospectus is a little more optimistic. Their projections, explained at the bottom of the linked page, have the Indians at 3% to win the division and 0.9% to win the wildcard.

"Wins in the Bank" and the Gambler's Fallacy

Several sites also post preseason standings projections. This is where the "gambler's fallacy" comes in. BaseballProjection.com has the Indians going 81-81 on the season, yet they have gone 11-18 so far.

Someone following the gambler's fallacy would think that the Tribe will go 70-63 the rest of the season to finish 81-81. In fact, Chone's projection should be thought of as a winning percentage, not a definite number of wins and losses.

That is, we should accept the fact that the Indians are 11-18, and assume they will play .500 ball (81 wins/162 games) the rest of the way. That means the Indians should expect to win 66-67 of their remaining 133 games

Pitching

BABIP

One of the most consistent statistics in baseball is a pitcher's batting average on balls in play (BABIP). The more a pitcher throws, the more his BABIP will settle in the .270-.290 range. This is true for nearly all pitchers, with only rare exceptions like knuckleballer Tim Wakefield and control specialist Greg Maddux.

That being said, if you see a pitcher with a high BABIP early in the season, you can expect that BABIP - and his overall performance - to improve. Conversely, pitchers with a low BABIP can expect an increase.

According to Fangraphs, David Huff, Jensen Lewis, and Aaron Laffey are already in the .270-.290 range. That's good news for all three, as they have pitched fairly well this season (outside of Huff's win-loss record) and should expect that performance to continue.

Cleveland's two best starters, Fausto Carmona and Mitch Talbot, have been beneficiaries of BABIPs in the .250s and can look forward to some minor regression. The same goes for relievers Chris Perez and Tony Sipp. Meanwhile, Joe Smith and Jamey Wright should see their already high ERAs rise with their BABIPs.

Meanwhile, Jake Westbrook and Justin Masterson should see their traditional numbers improve as their BABIPs decrease. As should Rafael Perez and Kerry Wood, which is good news for their respective ERAs. Hector Ambriz, pitching well by traditional measures, should only get better as his BABIP lowers. Not bad for a Rule V pickup.

Fielding Independent Pitching

While BABIP looks only at balls put into play, fielding independent pitching (FIP) does the opposite. FIP looks at walks, strikeouts, and home runs - the three things that are supposedly the pitcher's responsibility. In other words, FIP looks at those things that can't be affected by the quality of the defense.

FIP is calculated on the same scale as ERA, so the two can be compared easily. If a pitcher's FIP is lower than his ERA, he can expect improvement, and vice-versa. Of course the major caveat here is that if a pitcher plays with the same defense all year, how much can it really improve?

The bad news is that everyone except Jamey Wright, Justin Masterson, Rafael Perez, and Kerry Wood are posting FIPs lower than their ERAs. The even worse news is that Wright's FIP is only 0.40 points lower, and Wood's is based on only one inning of work.

Hitting

Lineup Analysis

BABIP can be used to study a hitter's luck. But unlike pitchers BABIP, hitter's BABIP normalizes to the hitter's past performance, not an overall league average. I'll leave that as an exercise to the reader.

Instead, for hitters we'll again refer to Dave Pinto's lineup analysis tool. Plugging in the top nine Indians in terms of OPS, here are the results. The lineup is certainly plausible in terms of defensive alignment and batting order, assuming Russell Branyan can still play some outfield.

This theoretical lineup would score 5.232 runs per game, much better than the 3.8 the team is currently scoring. That's an extra 232 runs for the year, or 23 wins. Of course, that's based on some small sample sizes, even for the regulars.

Friday, April 09, 2010

Sabermetric Bracketology: 2010 Post Mortem

Since I didn't get a chance to preview my brackets this year, here's a wrapup of how I did.

Straightforward Scoring: The "Pomeroy Bracket"

Again I used the Pomeroy Ratings for my straightforward bracket, and this year it was a success. Well, at least it was successful as getting three Elite 8 and one Final Four team can be, since that one Final Four team was Duke, who I also had winning the championship.

I can also thank Pomeroy for correctly putting Tennessee, Butler, and Xavier in the Sweet 16, and Baylor in the Elite 8. Outside of that, though, I mostly have to attribute my success to luck.

Upset Scoring: The Expected Value Method

I had one my upset scoring bracket two years in a row, but this year my luck ran out. And yes, it was luck, since part of my method involves picking which regionals to which I would apply my expected value methods (detailed here and here).

Those methods served me well in getting Cornell into the Sweet 16. But my major setback was a combination of bad luck and my own fault. My method calls plays the odds by putting all four number ones in the Final Four, and this year that wasn't a good strategy. On top of that, I decided to hedge my bets by picking Kansas over Kentucky in the finals, instead of relying on Pomeroy to have Duke as my champion again.

So, I finished in fourth, just out of medal contention. For next year, I'm going to study my previous brackets to see if it would have been better to use the expected value method on all four regionals, or whether I should continue to use "control" regionals. Stay tuned for those results.

Wednesday, February 24, 2010

Cleveland Indians: CHONE's WAR Gives Hope for a Winning Season

The Cleveland Indians may be rebuilding again, but it's hard not to be optimistic when CHONE predicts the team to win at least 81 games and finish second in the division.

Let's take a closer look at those projections, using WAR - Wins Above Replacement.

For those unfamiliar with the stat, WAR compares a given player to a replacement player - basically, any random guy you'd pull out of AAA. WAR is nice because it takes into account defense and position difficulty (it's harder to play shortstop than left field) in addition to offense.

It's said that a team of nothing but replacement players will win 47 games. So let's start with that as our baseline.

First, we'll look at the position players. FanGraphs includes this year's CHONE projections, and even gives a WAR number. So that makes things easy. Here are the position players on the 40-man roster.
PlayerWAR
Grady Sizemore5
Shin-Soo Choo3.1
Asdrubal Cabrera2.8
Jhonny Peralta2.2
Luis Valbuena1.5
Matt LaPorta1.3
Louis Marson1.2
Travis Hafner1.2
Brian Bixler1
Michael Brantley1
Mike Redmond0.8
Trevor Crowe0.8
Wyatt Toregas0.8
Jordan Brown0.6
Carlos Santana0.4
Andy Marte0.3
Nick Weglarz-0.1
Chris Gimenez-0.5
Jason Donald-1
Wes Hodges-1.1
Carlos Rivero-1.3

Sizemore and Choo are really good, but you already knew that.

Next, the Non-Roster Invitees.
PlayerWAR
Shelley Duncan1.5
Russell Branyan1.2
Brian Buscher1
Austin Kearns0.7
Luis Rodriguez0
Mark Grudzielanek-0.3
Damaso Espino-0.7
Niuman Romero-1.1
Beau Mills-1.2
Lonnie Chisenhall-1.8


Assuming we assemble a 25-man roster with 13 position players and 12 pitchers, here are the best possible position players we could take.
PlayerWAR
Grady Sizemore5
Shin-Soo Choo3.1
Asdrubal Cabrera2.8
Jhonny Peralta2.2
Luis Valbuena1.5
Shelley Duncan1.5
Matt LaPorta1.3
Louis Marson1.2
Travis Hafner1.2
Russell Branyan1.2
Brian Bixler1
Michael Brantley1
Brian Buscher1

OK, so we're short a catcher - that won't work.

Now, here's my best guess at what the actual opening day lineup will look like.
PlayerWAR
Louis Marson1.2
Matt LaPorta1.3
Luis Valbuena1.5
Jhonny Peralta2.2
Asdrubal Cabrera2.8
Michael Brantley1
Grady Sizemore5
Shin-Soo Choo3.1
Travis Hafner1.2
Mike Redmond0.8
Brian Bixler1
Andy Marte0.3
Russell Branyan1.2

The realistic lineup gives us 22.6 WAR, and the best-case gives us 24. Add that to our baseline of 47 wins, and we're already on pace to win 69-71 games.

Now, the pitchers. CHONE's website itself lists Runs Vs. Replacement for pitchers, which is nice. The common standard is that 10 runs = 1 win, so we'll divide the R vs. Rep by 10 to get each player's WAR.

Here are the pitchers on the 40-man.
PlayerWAR
Justin Masterson2.7
Fausto Carmona2
Aaron Laffey1.9
Jeremy Sowers1.5
David Huff1.4
Hector Rondon1
Jake Westbrook1
Hector Ambriz0.8
Carlos Carrasco0.6
Kerry Wood0.6
Mitch Talbot0.6
Rafael Perez0.6
Chris Perez0.5
Jensen Lewis0.5
Jess Todd0.4
Tony Sipp0.4
Joe Smith0.3
Jeanmar Gomez-0.8

CHONE didn't have a projection for Kelvin De La Cruz. But boy, it sure is high on Justin Masterson, isn't it?

The non-roster invitees.
PlayerWAR
Anthony Reyes0.9
Jason Grilli0.6
Mike Gosling0.2
Saul Rivera0.2
Frank Herrmann-0.1
Josh Judy-0.1
Zach Putnam-0.3

Alex White and Yohan Pino didn't get CHONE projections.

Here's the best possible group of 12 pitchers.
PlayerWAR
Fausto Carmona2
Aaron Laffey1.9
Jeremy Sowers1.5
David Huff1.4
Hector Rondon1
Jake Westbrook1
Anthony Reyes0.9
Hector Ambriz0.8
Carlos Carrasco0.6
Kerry Wood0.6
Mitch Talbot0.6

Rafael Perez is also at 0.6, so feel free to substitute him in for Carrasco, Wood, or Talbot as you see fit.

Here's my best guess at the opening day staff. I wasn't sure quite how it was going to turn out, so I started with the guys under contract (Westrbook, Wood, Carmona, and Perez), added rule 5 pickup Ambriz, then just went in descending order by WAR after that.
PlayerWAR
Jake Westbrook1
Kerry Wood0.6
Fausto Carmona2
Rafael Perez0.6
Hector Ambriz0.8
Justin Masterson2.7
Aaron Laffey1.9
Jeremy Sowers1.5
David Huff1.4
Hector Rondon1
Carlos Carrasco0.6
Mitch Talbot0.6


The best possible scenario clocks in at 15 total WAR, and the "realistic" one comes in just under that at 14.7. Add that to our 47-win baseline and the 22-24 wins by the hitters, and we're looking at a team that should win at least 80 games and could win as many as 87.

Of course, a word of caution: it was these types of sabermetric projections that predicted the Tribe to win the Central in '06, '08, and '09. But still, it's February, so why not be optimistic?

Thursday, April 30, 2009

Jhonny Peralta: Cold Like the Weather

From Andy Castrovince's latest mailbag:
I'd be interested to see Jhonny Peralta's batting average in warm weather vs. cold weather. He was red hot down in Arizona this spring, but now he seems ice cold, much like the first six weeks of last season. Great Odin's raven, am I on to something here?
-- Tim R., San Diego

I'd love to help you find those numbers, but, much like the translation of the name for your hometown, scholars maintain that the ability to calculate such statistics was lost hundreds of years ago.

Actually, I asked hitting coach Derek Shelton and media relations director Bart Swain if they've ever heard of such a stat, and neither has. You'd really have to be the obsessive-compulsive type (even by baseball statistician standards) to calculate those numbers, especially when you consider the temperature at first pitch can take a drastic dip by the last pitch.

Well, when it comes to the Indians (and especially Peralta, apparently), I am that obsessive-compulsive baseball statistician. After reading the question, I immediately thought, "I bet Retrosheet tracks weather information. Sure enough, they do.

If you want me to bore you with the details, email me. Otherwise, let's skip to the pretty picture:



The short explanation is that I took Peralta's batting average (hits/at bats, obviously) for each gametime temperature reading. As you can see, there is an upward trend. Thoughts:
  • There's no context. Maybe all, or at least most, hitters follow the same trend. That would make this "revelation" meaningless.

  • Correlation is not causation. It's warm at midseason, when hitters are thought to be at their best, and when the ball is said to travel better. It's cold at the beginning of the season, when hitters aren't yet in "midseason form", and at the end, when fatigue starts to set in. So, again, maybe this is a trend for everyone and not just Jhonny.

  • I'd be interested to see weather-related trends for other statistics, starting with BABIP and OPS, and perhaps moving onto the more advanced stuff.

Saturday, March 21, 2009

Stathead Bracketology

I participate in two NCAA bracket competitions every year. One (through this site) uses fairly straightforward scoring based on the round and nothing else. The other uses upset bonus points: when the lower seed wins, you get the normal points for the win (one for the first round, two for the second, etc.) plus bonus points worth the difference in the two seeds. For example, for picking a 12-seed over a 5-seed in the first round, you get one point for the win plus 12-5=7 bonus points, for eight points total.

Now, being statistically inclined, I wanted to use mathematical methods in each. Here's what I did.

Straightforward Scoring: The Pomeroy Ratings

For my straightforward bracket, I decided to employ the Pomeroy Ratings. I used the Fremeau Efficiency Index Ratings with great success in our BCS Bowl Pick 'Em a few months ago, so I wanted to go with a similarly statistically-inclined system for the basketball pool.

I did it the simplest way possible, too. For each matchup, first round through final, I picked the team with the higher Pomeroy Rating. (Note that the ratings are being continuously updated through the tournament, so the ratings as they are right now do not match what they were when I made my picks the day before the tournament started.) Some notes:
  • Memphis is number one in the ratings, and therefore my champion.

  • The ratings "predicted" Wisconsin's win over Florida State.

  • All of the 6 seeds were apparently undervalued by the selection committee. Going by the Pomeroy Ratings, all of the sixes except Marquette were expected to beat the 3 seeds in their respective brackets and advance to the Sweet 16.

  • West Virginia was in Pomeroy's top 10, which should have gotten them into Elite 8. But I guess Dayton had something to say about that.

  • I did "cheat" and pick Cleveland State in the first round, since my dad was a two-time basketball letterman there. I'd be kicking myself this morning if I hadn't picked them.


My "straightforward bracket" can be found as a Google Doc here.

Upset Bracket: The Expected Value Method

Last year, I outlined a method for using expected values to make picks in a pool with upset points. Well, I'm happy to say that my method worked. Thanks to a bevy of first round upsets last year, I built a big lead and was able to withstand a single competitor to win the pool. I did have to sweat out the last rounds as my bracket started to fall apart, but I was so far ahead that it didn't matter.

As with last year, I used a different method for each regional. Two were methods (b) and (c) from Part 2 of last year's post. The other two were "controls." On one regional, I would pick only the top seeds for every game. On the other, I would pick by feel. Actually, I even took the human element of "feel" out this year, instead picking by the Pomeroy method chosen above.

Thanks to the Pomeroy bracket I filled out first, I was able to pick and choose which method I would use on each regional. And I will admit that I did allow for the human element to creep in when making my decisions. I used the two different expected value upset methods on the Midwest and South regionals. This mainly allowed me to advance West Virginia to the Elite 8 again (whoops), and pick Cleveland State in the first round.

I made the East my Pomeroy regional, mainly so I could take Wisconsin in the first round and put 6 seed UCLA in the Sweet 16. That left the West as my higher seeds only regional. Unfortunately, that means I won't have Memphis as my winner in this bracket, but maybe hedging my bets isn't a bad thing.

Once I got to the Final Four, where the seeds don't matter anymore, I went back to the Pomeroy Ratings to determine the finals participants and eventual winners.

First Round Results

My straightforward bracket isn't doing so hot, garnering only 23 of a possible 32 points, with one Elite 8 team (West Virginia) and one Sweet 16 team (Utah) down.

My bonus points bracket is doing fantastic, though, as it nailed all three 12 seed wins plus the Cleveland State upset. I am down an Elite 8 team in West Virginia, but the 57 points I did pick up should again give me a lead that's difficult to catch.

Wednesday, January 07, 2009

Cleveland Indians WAR Spreadsheet

Thanks to the hard-working Sky Kalkman over at Beyond the Boxscore, here's a spreadsheet of the Indians 2009 predicted Wins Above Replacement (WAR).



Some notes:
  • For hitting and pitching, I used the CHONE projections. On Sky's suggestion, I also used the quick-and-dirty formula (OBP*1.75 + SLG)/3 to approximate wOBA.

  • I included Luis Valbuena on offense to get closer to the "recommended" number of outs. As you can see, I'm still short. I'm not doing this as an endorsement of Valbuena over Barfield; CHONE has nearly identical plate appearances and wOBA for the two, making them interchangeable for the purposes of this exercise. I did leave out Andy Marte, since the popular assumption is that his time in Cleveland is short.

  • While my spreadsheet falls short of the recommended outs on offense, adding players (and therefore outs) actually increases offensive WAR in most cases. So, take this as a low-end projection.

  • Baserunning numbers are from Baseball Prospectus's EqBRR. I just used the 2008 numbers, so obviously there's room for improvement there.

  • For fielding, I used UZR from Fangraphs. I took a straight average of each player's UZR from 2006-2008 (when available). Obviously, a weighted average probably would have been better.

  • I also ignored position when taking UZR. This is mostly because I'm convinced the Indians will try some experimenting on the infield to get it right. But at the same time, I didn't want to guess how many innings each player would log at each position.

  • For the rotation, I started with the four known quantities - Lee, Carmona, Pavano (who will hopefully top that inning projection), and Reyes. Then I started filling in the best available players based on the CHONES until I got to the recommended inning count.

  • For the relievers, I followed the current Indians.com depth chart until I got to the recommended number of innings. But that doesn't necessarily mean I endorse that depth chart.


What does this all mean? If I filled in the spreadsheet correctly, the Indians project to a 90 win team. That's right on the cusp of the playoffs, which is exactly where the Indians want to be.

Sunday, December 21, 2008

Notre Dame-Hawaii By The Numbers

Vegas Odds

The line is starting to move on this one. Originally Hawaii was a 1.5 point favorite, but now they're anywhere between there and a 1.5 point underdog to the Irish. They're a saying that you automatically give the home team 3 points. Well, despite the fact that Notre Dame will be wearing the dark jerseys, the Warriors have the home field advantage in this one. Of course, I've also heard that Notre Dame automatically gets 3 points in their favor on the line due to their following. Those 3 points of course cancel out the three given to Hawaii for home field advantage.

Computer Rankings

Below are the six computer rankings used by the BCS, and Notre Dame and Hawaii's position in each. As you can see, it's very tight.
55
Notre DameHawaii
Sagarin6061
Anderson & Hester5763
Billingsley6965
Colley6059
Massey59
Wolfe6259
Average60.561

But those aren't the only computer rankings out there. Jeff Sagarin also includes on his website (rather bitterly) another ranking, called his PREDICTOR, that relies heavily on margin of victory. In that, Notre Dame is 62nd and Hawaii 95th.

The Fremeau Efficiency Index, used by Football Outsiders and housed at BCS Toys, has Notre Dame at 47 and Hawaii at 101. In addition, BCF Toys is calling the game a lock, with an 88% degree of confidence that Notre Dame will win 30-10.

Stastical Trends

Finally, here's a quick-and-dirty regression analysis. For each game this season, I took Notre Dame's points for and against, and rushing, passing, and total yards for and against. Excel gave me the following coefficients:
StatCoefficient
Y-Intercept0.76
Points For0.01
Points Against-0.03
Rushing For0.06
Rushing Against-0.07
Passing For0.06
Passing Against-0.07
Total For-0.06
Total Against0.07

Like I said, it's a very quick-and-dirty model. I'm not even sure if the variables I used are reliable predictors of wins and losses. Of course, the small sample size of 12 games is also an issue. But this is all something I can work on next year.

Anyways, plugging Hawaii's season averages into this model gives Notre Dame a 57% chance of winning.

Conclusion

If you're a fan of what the stat nerds have to say, this will be a close one, with a slight edge to Notre Dame. Still, there are many more factors to consider, and time permitting I'll cover those in my next preview. Go Irish!

Wednesday, December 10, 2008

Cleveland Indians Sabermetrics 101: DIPS and RAR

In the last installment, we discovered that all pitchers' BABIP tend to fall in the .290 to .310. It follows then that if we want to gauge a pitcher's talent level, we need to look at everything besides balls hit into play. "Everything else" consists of strikeouts, walks, and homeruns, generally considered the "three true outcomes." There are a number of stats that study the three true outcomes, and together they are referred to as defense-independent pitching statistics - DIPS.

Before I go into examples of DIPS, I need to define a few terms.

replacement level: This is a popular concept among statheads. A replacement level player is one that is easily available as a mid-season free agent signing or a AAA call-up. Indians fans saw many replacement level players make starts for the Indians last year, Matt Ginter for example.

Runs Above Replacement, RAR: If "replacement level" is the amount of production you can get out of a player off the scrap heap, you would expect your regular players to be able to perform above that level. For pitchers, this means comparing the number of runs a pitcher gave up over a certain number of innings and comparing it to the number of runs a replacement player would have given up over the same number of innings. This is a tally of runs "saved" compared to a replacement pitcher, so a high positive number is better.

RAR for pitchers is calculated by taking the replacement level ERA, which is generally taken to be 5.75, subtracting the player's ERA, diving by 9 (since ERA is a measure of runs given up per 9 innings), and multiplied by innings pitched:

(5.75 - ERA) / 9 * IP

Wins Above Replacement, WAR: It's generally accepted in sabermetric circles that 10 runs is equal to 1 win. So, to find out how many Wins Above Replacement a pitcher earned, their RAR is divided by 10.

DIPS

Beyond the Box Score's article on this topic covers many pitcher stats. I'll let you look through those at your own leisure. The most advanced is a new stat called tERA. It takes the three true outcomes mentioned above, plus HBP and percentage of hits that were ground balls, line drives, infield flies, and outfield flies, plus takes the ballpark into factor. With input that complicated, it has to be accurate, right? Well, you and I can just take their word for it for the time being.

Since the goal for many sabermetricians is to take luck out of the equation, StatCorner - the same people that created tERA - also created xIP, expected Innings Pitched. Put simply, xIP tries to determine what each play "should have been" (ie, a screaming liner that was caught is changed to a hit, and a blooper that dropped is changed to an out) to give a more accurate look at the pitcher's workload.

Indians RAR/WAR in 2008

How did the Indians fare in 2008?
PitcherxIPtERARARWAR
Cliff Lee2222.64778
CC Sabathia122.673.26343
Fausto Carmona1214.64151
Zach Jackson573.81121
Aaron Laffey92.674.77101
Paul Byrd1285.4050
Jake Westbrook334.4050
Matt Ginter21.333.7650
Anthony Reyes324.7440
Scott Lewis234.5230
Jeremy Sowers1185.89-20
Bryan Bullington9.677.49-20
Tom Mastny1.6723.83-30
Total Result98279.13162.4516.24

As you can see, by this methodology Cliff Lee won eight games all by himself. Meanwhile, Matt Ginter, Bryan Bullington, and Tom Mastny were almost the definition of replacement level.

But that 162.45 total RAR means nothing without context. Cleveland finished fifth in the AL in RAR in 2008, just behind Boston and Tampa Bay, and just ahead of Anaheim and Minnesota. Toronto and the White Sox were almost 50 points ahead of their closest competitors, thanks to aces (Roy Halladay and Mark Buehrle) that scored in the 70s with a solid supporting staff. (Halladay and AJ Burnett are worth stars, but the rest of the Toronto rotation is very underrated.

Indians RAR/WAR in 2009

So, how do the Indians look in 2009? Will they need to add another starter?

The calculations for tERA and xIP are beyond my abilities at this point, so I cheated and used the innings pitched and ERA predictions from the 2009 Marcels. Here's what Marcel have to say for the guys currently on the 40 man roster.
PitcherIPERARARWAR
Cliff Lee1803.80394
Fausto Carmona1354.13242
Zach Jackson834.7291
Aaron Laffey1124.26192
Jake Westbrook934.16162
Anthony Reyes844.45121
Scott Lewis724.00141
Jeremy Sowers1274.82131
Total Result88634.34147.0314.7

Uh oh. That 147.03 RAR would be, by 2008 standards, ninth best in the AL. But there are a few things to remember. Marcel is admittedly a "dumb" system, and it looks at three years of data. That means it's taking Cliff Lee's disappointing 2007 and Fausto Carmona's nightmare 2006 into account. Also, RAR is dependent on innings pitched. So, once the Indians settle on their five best starters, and give them the innings that went to "experiments" last year, the numbers should improve.

Wednesday, December 03, 2008

Cleveland Indians Sabermetrics 101: BABIP

Inspired by this post at Beyond the Boxscore, here's my first "homework assignment" for Saber-Friendly Blogging 101.

Batting Average on Balls In Play, BABIP, is essentially batting average for everything except strikeouts and walks. I'll let the article above, Wikipedia, and the Sabermetric Wiki give you the details.

BABIP for hitters relies on many factors and will vary from player to player. But for pitchers, BABIP always seems to fall in the .290 to .310 range. What does this mean? If a pitcher is widely outside of that range one year, you can expect them to regress back to those numbers the following year, and their overall performance should follow. (I'm sure there are cases of certain pitchers having consistently high or low BABIP numbers, but I don't know of any offhand.)

So, how did Indians pitchers fare in 2008? To find out, you can check The Hardball Times or Fangraphs. Or, if you'd rather do the work yourself (or, like me, didn't find out about the THT and Fangraphs page until after doing the work), you need to start with their 2008 stats.

If you're not up to the task of setting up a MySQL database of baseball stats, you can simply copy and paste them from Baseball-Reference's 2008 Indians page into Excel. To determine At Bats, I took BFP (Batters Faced by Pitcher) minus Bases on Balls and Hit By Pitch. (I assumed IBB totals were already included in IBB.) I also ignored Sacrifice Flies because that information wasn't readily available. My results were almost identical to those from the Hardball Times. Fangraphs' were a little different, but as the Sabermetric Wiki mentions, there are several variations to the formula.

PlayerABBABIP ERA
Rich Rundles19.3851.80
Tom Mastny89.34410.80
Edward Mujica157.3286.75
Rafael Betancourt284.3115.07
Zach Jackson221.3105.60
Masahide Kobayashi229.3064.53
Rafael Perez288.3043.54
Cliff Lee852.3012.54
Jeremy Sowers491.3015.58
Jensen Lewis260.3003.82
Aaron Laffey369.2944.23
Fausto Carmona470.2945.44
Jake Westbrook131.2623.12
Anthony Reyes129.2591.83
Scott Lewis91.2222.63
Jonathan Meloan5.0000.00


I included At Bats in this table to illustrate a point: in general, as at bats increased, the pitcher's numbers moved more to the .290-.310 range. Rich Rundles and Jon Meloan only faced a handful of batters, so their numbers can largely be ignored. But you could almost argue the same for Scott Lewis and Tom Mastny.

Now, this table is good news for Tom Mastny and Ed Mujica, and even Rafael Betancourt and Zach Jackson to some extent. All posted high ERA and high BABIP. But their BABIP should go down in 2009, and their other stats should improve as a result. Conversely, Anthony Reyes and Scott Lewis will probably see their spectacular 2008 numbers fall back to earth. Jake Westbrook will probably see a decline as well, once he's finally healthy.

If BABIP holds true, that middle group should stay about the same. That's great news for Cliff Lee, Rafael Perez, and Jensen Lewis, as well as Aaron Laffey and Masahide Kobayashi to some extent. But it's also bad news for Jeremy Sowers and Fausto Carmona.

But BABIP is by no means a be-all, end-all predictor. For example, while Rafael Betancourt was on the edge of the expected BABIP range, his ERA was abnormally high (for him) due to a lingering injury that kept him from throwing his fastball, which is his best pitch. And while Cliff Lee was right in the middle in terms of BABIP, he'll still be hard-pressed to repeat the phenomenal year he had in 2008. Still, his BABIP numbers do show that 2008 wasn't entirely luck, and that Lee should still do very well in 2009.

Sunday, November 30, 2008

Cleveland Indians: Is Being Average An Improvement?

Recently, Walk Like a Sabermetrician took a look at 2008 hitting by position. They calculated runs created and runs created per game for each position, and used that to create a Runs Above Average metric that compares each player to the average production at their position.

My hypothesis has always been that you should want at least league-average production at every position - or at least a net of league average production.

So how did the Indians fare in 2008? Not too well. If you scroll down to the second to the last table in the Walk post, you see that the Indians were great at shortstop and center field, very good at catcher, and varying degrees of bad everywhere else. That's only three of nine positions above average, and a net of -11 runs above average.

The next step is to predict how the Indians will fare in 2009, to see what, if any, changes need to be made. I couldn't figure out how RAA was calculated. However, I did take the 2009 Marcels and calculated runs created and RC/27 for each of the Cleveland Indians key figures. Having that, I could compare that to the positional averages from the Walk article.

Now, I will admit there are two problems with my methodologies:
1) I'm using the generic RC/27, while "Walk" is using RC/G. RC/G uses the average number of outs per game (since the home team doesn't bat in the ninth when they're leading), which from what I could tell is normally around 25.5.
2) I'm comparing 2009 projections (from an admittedly simplistic projection model) to 2008 numbers, when ideally I should be comparing to several years worth of data.

That being said, here are Cleveland's RC/27 numbers for 2009, according to Marcel:
Choo6.63
Sizemore6.5
Hafner5.91
Martinez5.54
Shoppach5.36
Garko5.21
Francisco5.12
Peralta5.1
Cabrera4.76
Gutierrez4.48
Dellucci4.36
Carroll4.06
Barfield3.91
Marte3.7

Unfortunately, the Marcels don't include lefty/righty splits, but that would be a nice thing to look at in the future. Keeping that in mind, the best lineup by these numbers would have Kelly Shoppach catching; Victor Martinez at first; Asdrubal Cabrera, Jamey Carroll, and Jhonny Peralta in the infield (I'll leave the position argument for another article); Ben Francisco, Grady Sizemore, and Shin-Soo Choo in the outfield; and Travis Hafner at DH.

The 2009 projections give Cleveland four players above the league average for their position - Shoppach at catcher, Peralta at short, Sizemore in center, Choo in right, and Hafner at DH - with Cabrera around the league average at second and Martinez not too far behind the league average at first. (Or, if you prefer, you can call Peralta a league-average third baseman and Cabrera an above average shortstop.) You can increase that number to five if you count that Francisco's numbers are above-average for a center fielder and Sizemore's are still above average for a center fielder.

Fans of the Indians should see several nice things in these numbers: Martinez' ability to handle first without much of an offensive number, Cabrera's bat catching up to his glove, Peralta able to hit enough to be at least an average first baseman, and Hafner bouncing back to be productive again. Of course, these are just projections, but the offseason is always a time for optimism.

Now, back to the numbers. To recap, the Indians have five above-average positions (comparing Francisco to a center fielder), one average, one slightly below average, and one well below average. Working solely off these numbers, and thinking only about 2009, the wise choice would be to trade Kelly Shoppach for a league-average infielder. In that situation, Martinez would move back to catcher and Garko would enter the lineup at first base.

Thanks to Peralta and Cabrera's versatility, that infielder can be from any position, although a 5.1 RC/G third baseman would obviously be more productive than a 4.8 RC/G second baseman or 4.4 RC/G shortstop. The increased production from Carroll to the new infielder would more than compensate for the dropoff from Shoppach's 5.36 RC/27 to Garko's 5.21.

Of course, I am in no way suggesting a trade based solely on these findings. But what I am suggesting is that, by virtue of having more players at or above their position's averages in 2009, the Indians can look forward to a better season offensively.

Saturday, October 18, 2008

2009 Cleveland Indians: A Modest Proposal

ChooRF
SizemoreCF
Garko1B
HafnerDH
Martinez3B
ShoppachC
FranciscoLF
Peralta2B
CabreraSS

How's that for a lineup in 2009? It just might work.

Offense

One way to construct a good lineup is to put above average hitters at every position. In the below table, each player is followed by two numbers. The first is their career OPS. Eventually I'll edit this to use their projected 2009 OPS, but for now I'll use their career totals.

The second number is the average OPS among qualified batter at that position in 2008. For example, in 2008 nine catchers had enough at bats to qualify for the batting title, and those nine catchers had an average OPS of .776.

NamePosC OPSP OPS
ShoppachC.776.794
Garko1B.800.776
Peralta2B.771.787
Martinez3B.832.824
CabreraSS.732.764
FranciscoLF.770.858
SizemoreCF.861.766
ChooRF.870.816
HafnerDH.924.895

The first thing that probably jumps out to Cleveland fans is the heavy assumption that Ryan Garko and Travis Hafner progress back to their career averages. If they do, though, they'll both be above average players at their positions.

The second thing that jumps out is that first basemen had the same average OPS as catchers, and were below second basemen. That's a statistical anomaly I should look into if I want to improve this study.

The third thing that should jump out is Jhonny Peralta's OPS compared to each infield position. Peralta is slightly above average as a shortstop, slightly below average as a second baseman (by 2008 standards, at least), and would be eaten alive as a third baseman. Victor Martinez, meanwhile, is already above average as a third baseman, and that doesn't count the mythical batting improvement of a catcher moving out into the field.

In the outfielders, many fans want the Indians to sign a hard-hitting corner outfielder. But that may not be needed. Both Shin-Soo Choo and Grady Sizemore hit better than the average left fielder - the strongest outfield position. That should be enough to outweigh Ben Francisco's center field-like hitting.

Combining all the positions above, you get an average-league-average OPS of .807, while the Indians give an average OPS of .817. Granted, that's not weighing everyone's plate appearances individually, but the basic numbers show good things for the Tribe.

Defense

The big thing here, of course, is how Victor Martinez would handle third base. He was signed as a shortstop, but played catcher his entire minor and major league career.

Jhonny Peralta's defense should improve by moving to an easier position. The team's defense as a whole should improve by having better defensive players at shortstop (Cabrera or Peralta) and catcher (Shoppach over Martinez).

Order

I plugged these players into David Pinto's lineup analysis tool, and the results can be found here. The lineup listed at the top of this article isn't the best one the tool returned, but it's the one that makes the most sense. (Martinez and Hafner in the top two spots could get anyone fired.)

The only concern with the lineup is lefty-righty matchups. Five out of six consecutive hitters (counting Cabrera) are lefties. But they are broken up by Garko, who does much better against lefties than righties.

Conclusion

If the Indians want to get back to the playoffs, the parts may already be there. Sure, moving Martinez to third and Peralta to second is far-fetched. But am I any crazier than the professional sportswriters who think all the Indians' problems can be solved by signing Manny Ramirez, a guy who burned all his bridges in Cleveland long ago?

As I mentioned above, this is still a work in progress. I welcome any input on my methods or my conclusions.

References

ESPN's MLB stat page for the breakdown of qualified hitters by position
Fangraphs, for players' career OPS numbers

Saturday, July 26, 2008

Jeff Samardzija PITCHf/x

For those who missed it, Jeff Samardzija made his major league debut for the Cubs against the Marlins yesterday, giving up 1 run in two innings of relief work. Here's a PITCHf/x study of how The Shark fared. This is my first time working with PITCHf/x, so any and all feedback is welcome. I originally was going to work with the data myself, but then I ran across BrooksBaseball.net, which does all the dirty work for you. I promise I'll learn how to do everything from the ground up one day, but just not today.

You can see all the charts and data Brooks Baseball provides here. I'll go over the ones that are easier to understand (for me, at least) below. For all images, click for a larger version.


This chart gives the speed of each of Samardzija's pitches. As you can probably see, he started off almost exclusively with fastballs to get himself adjusted. Then, he alternated fastballs and breaking pitches (sliders and changeups, according the MLB Gameday) over the rest of the appearance.

He was in the 96-98 MPH range with his fastballs, topping out at 99. His breaking balls were all right around 84 MPH. If he can keep that 10+ MPH difference between his fastball and his changeup, he'll be quite effective.


Next, here is a plot of balls and stikes. The blue points, labeled "X", are balls put in play.

As you can see, Samardzija pounded the zone pretty well for most of the outing. The extreme outlier was a pitchout/wild pitch. This graph is from the catcher's point of view, meaning that that pitch was high and outside to the left-handed batter.


Here's a graph of Samardzija's release point from pitch to pitch. Having a consistent release point is important for all pitches. If Samardzija had a different release point for his fastball and changeup, for example, batters could pick this up and know when each pitch is coming. Since the main purpose of a changeup is to look like a fastball, this would completely destroy the effectiveness of the pitch.

Samardzija had a consistent release point across all three pitches, which is a plus. The one outlier, I believe, is the pitchout - so absolutely nothing went right there.


This one shows the speed of pitches color coded by result - ball, strike, or in play. The gap in data represents the bottom of the seventh, when the Cubs were at bat. This chart more clearly shows how Samardzija started with just fastballs, then mixed it up, then went mostly offspeed in the eighth.



I'm not sure how much information you can get out of these last two (it's possible that someone can't, but I can't). They are interesting to look at, however. These graphs show how, on average, each type of pitch breaks. The first is an overhead view, as if you were in the blimp with Samardzija at the top of the graph and home plate at the bottom. The second is a side view with the pitcher on the right and home at the left (in other words, as if you were sitting on the first base side).

What can be learned from this? Well, I'm sure quite a bit, but here's what my simple mind picked up: First, Samardzija pounded the zone, showing he wasn't afraid to throw strikes. Second, he had a consistent release point for all pitches, which is a positive. Third, knowing he was only going to throw a few innings (he was a starter in the minors), he was able to max out and throw his fastball in the high 90s, another positive. Finally, his breaking stuff was 10+ MPH slower than his heater, meaning that those pitches will be very effective at keeping batters off balance.

Sunday, July 20, 2008

Cleveland Indians: Quick & Dirty Second Half Projections

The Hardball Times recently released a tool that uses the Marcels projection system to create quick-and-dirty projections of a player's performance over the rest of 2008. Picking up on what's already been done for the Mets and Brewers, here's what the tool had to say about the Indians.

Hitters

The Marcels are best for rate stats, such as batting average, OBP, SLG, and OPS. Here, I measure improvement or regression in terms of the difference between a player's OPS-to-date and their "balance" OPS, or their projected OPS from now until the end of the year. This is labeled as OPS_delta.

Massive Improvement

(> 3 Standard Deviations)
Josh Barfield
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
2008366000000000000
2008 (proj balance)66131123337127200.2640.3080.3874810.6950.695
2008 (proj total)69137129337127200.2520.2940.3694810.663


This one really isn't fair, since Barfield will obviously improve on his 0-for-6 start. Unfortunately, .695 and .663 are still below league-average OPSs.

Much Improvement

(> 1 Standard Deviation)
Jorge Velandia
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
20081016153100140.20.250.267400.517
2008 (proj balance)66105942550310180.2660.3420.424010.7620.245
2008 (proj total)761211092860311220.2570.3290.3994410.729


Travis Hafner
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
2008461881573490423440.2170.3320.355540.682
2008 (proj balance)66268226611411238440.2710.3840.49711230.8820.2
2008 (proj total)112456383952311661880.2490.360.43716770.797


I can't guarantee all the Jorge Velandia fans out there that he'll play 66 games over the rest of the season. But I can back Marcel's other guarantee - that Travis Hafner, when healthy, will greatly improve over what he's done so far.

Some Improvement

(< 1 Standard Deviation)
Andy Marte
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
2008339082153014230.1830.230.2562110.486
2008 (proj balance)6617916438101413310.2310.2950.3766220.6720.186
2008 (proj total)9926924653131517540.2150.2710.3368330.607


Asdrubal Cabrera
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
2008531891613070122410.1860.2880.2484010.536
2008 (proj balance)6623420851111424360.2480.3310.3697720.70.164
2008 (proj total)11942336981181546770.2210.3090.31611730.625


Victor Martinez
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
20085421719855110016230.2780.3350.3336610.668
2008 (proj balance)6626323669150725290.2940.3680.44910620.8170.149
2008 (proj total)120480434124260741520.2860.3520.39617230.748


Franklin Gutierrez
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
20087623121446101310500.2150.2630.3136740.576
2008 (proj balance)6619918547101513380.2530.3070.47420.7070.131
2008 (proj total)14243039993202823880.2330.2820.35314160.635


Ryan Garko
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
20088432829170110725470.2410.320.35110290.671
2008 (proj balance)6625623063120819370.2730.3470.43410070.7810.11
2008 (proj total)1505845211332301544840.2550.330.387202160.718


Michael Aubrey
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
2008930275002330.1850.2670.4071100.674
2008 (proj balance)662181965091720300.2560.3310.4188220.7490.075
2008 (proj total)752482235591923330.2470.3230.4179320.74


David Dellucci
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
20087724822451130817530.2280.30.3938860.692
2008 (proj balance)6621118846102820390.2460.3290.4378230.7660.074
2008 (proj total)143459412972321637920.2360.3120.41317090.725


Shin-Soo Choo
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
2008371311112681316290.2340.3410.4054520.746
2008 (proj balance)6623220553132524420.260.3450.4198630.7640.018
2008 (proj total)10336331679213840710.2510.3420.41413150.756


There are positives and negatives to be taken away from this group. First, Shin-Soo Choo, who has been hitting well (especially compared to his teammates), is actually due for a minor improvement over the rest of the season. Second, Victor Martinez, Ryan Garko, and David Dellucci are all due to hit at an above-league-average clip in the second half. That's just what this team needs, and what those players need individually. Now the negatives: yes, Andy Marte, Asdrubal Cabrera, and Franklin Gutierrez are due for some improvement. But that still won't make their final numbers very pretty.

Exactly the Same


Jhonny Peralta
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
200888372345902511623720.2610.3070.47816500.785
2008 (proj balance)6627725168151924530.270.3370.44711310.7850
2008 (proj total)15464959615840225471250.2650.3180.46527810.783


Earlier this year, I noticed that Jhonny Peralta's career slash stats were almost exactly league-average. Peralta once again will live up to that Mr. Average billing; he'll see an increase in his OPB but a slight dropoff in his slugging percentage.

Some Regression

(< 1 Standard Deviation)
Kelly Shoppach
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
20085818616841120711610.2440.310.447450.75
2008 (proj balance)6621019348120715530.2510.3130.4268230.739-0.011
2008 (proj total)1243963618924014261140.2480.310.43315680.743


Jamey Carroll
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
2008682482135774020380.2680.350.3387270.688
2008 (proj balance)662392135492222310.2540.3340.3427330.676-0.012
2008 (proj total)134487426111166242690.2610.3360.34145100.676


Ben Francisco
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
20086426623970190822410.2930.3570.47311320.83
2008 (proj balance)6627224870171823420.2810.3460.45511320.801-0.029
2008 (proj total)1305384871403611645830.2870.350.46422640.814


Grady Sizemore
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
2008924263661002042353720.2730.3740.53819760.912
2008 (proj balance)66303264751631234490.2850.3770.50113350.878-0.034
2008 (proj total)15872963017536735871210.2780.3750.523330110.897


Casey Blake
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
200888345305862301030640.2820.360.45613970.815
2008 (proj balance)6625723260141821440.2610.3330.43210040.765-0.05
2008 (proj total)15460253714637118511080.2730.3460.446239110.792


Most Indians fans know that these are the guys who are having good years. Unfortunately, that party's about to end for some. The good news is that the dropoff is slight. For most guys, Marcel predicts most of the regression to come from slugging percentage, not on base. For example, Marcel says Grady Size is due for a 0.003 decline in OBP but a 0.037 drop in SLG. This is because Marcel is an imperfect model (no offense, of course, because it wasn't meant to be perfect) based on past performance. Grady's never hit 23 HR before the break before, and that's taken into account for his second-half performance. Meanwhile, those urging the Indians to sell high on Casey Blake may finally have their statistical proof.

Much Regression


Sal Fasano
YearGPAABH2B3BHRBBKBAOBPSLGTBHBPOPSOPS_delta
2008415135100250.3850.4670.462600.928
2008 (proj balance)6624622654111815540.2390.30.3968950.696-0.232
2008 (proj total)7026123959121817590.2470.310.3999550.709


Sorry Sal, but we knew you couldn't keep it up for that much longer.

Pitchers

The Marcel projections returned two useful rate stats - ERA and WHIP. I added these two together to create a Total_delta for each pitcher.

Massive Improvement

(> 3 Standard Deviations)
Brian Slocum
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
200820272810141526
2008 (proj balance)104.48132102.51001-22.52-1.5-24.02
2008 (proj total)3017.893113113.42527


Brian Slocum was the curve-wrecker on this one. Then again, 2 IP with 6 ER to date will do that for you.

Much Improvement

(>1 Standard Deviation)
Tom Mastny
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
20086113.018.3139802.5247312
2008 (proj balance)414.39686301.933213-8.62-0.59-9.21
2008 (proj total)1029.521421151102.2979415


Like Slocum, Tom Mastny's numbers were affected by his small sample size. (Unlike Slocum, Mastny also had that ugly emergency start to his credit.)

Some Improvement

(< 1 Standard Deviation)
Jeremy Sowers
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
2008997.5244.366261911.922141037
2008 (proj balance)664.733037191111.61145516-2.79-0.31-3.1
2008 (proj total)15156.3974103453021.793591553


Rafael Betancourt
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
200842064248421401.48185828
2008 (proj balance)2903.58292926911.32126411-2.42-0.16-2.58
2008 (proj total)7105.027177682311.413111239


Juan Rincon
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
20082406.112833201621.75133519
2008 (proj balance)1604.17192116811.579029-1.94-0.18-2.12
2008 (proj total)4005.324754362431.68223728


Paul Byrd
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
200818185.47102121441751.354352362
2008 (proj balance)12124.766982351531.42951137-0.710.05-0.66
2008 (proj total)30305.18171203793281.377303499


Jensen Lewis
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
20082404.4134.738261921.64160417
2008 (proj balance)1603.942425201011.46109210-0.47-0.18-0.65
2008 (proj total)4004.225863462931.57269627


Edward Mujica
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
20081304.8516.71512411.146939
2008 (proj balance)904.111118301.284715-0.750.14-0.61
2008 (proj total)2204.55282620711.19116414


Those who have watched Jeremy Sowers' last two starts know he's already started to improve. Meanwhile, it's nice to see Rafael Betancourt, Paul Byrd, and Jensen Lewis - the poster children for regression on the pitching staff in 2008 - at least have some room for improvement over the rest of the year.

Some Regression

(< 1 Standard Deviation)
Rafael Perez
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
20084303.1642.740431611.31177515
2008 (proj balance)2903.652926261011.241203120.49-0.070.42
2008 (proj total)7203.367266692621.28297827


Aaron Laffey
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
200815153.6189.791412991.343841036
2008 (proj balance)10104.216160372041.312616280.6-0.030.57
2008 (proj total)25253.851511517849131.336451664


Fausto Carmona
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
200810103.15854233831.59255120
2008 (proj balance)773.943939261621.381733170.84-0.210.63
2008 (proj total)17173.449793495451.5428437


Jake Westbrook
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
2008553.1134.73319711.15139512
2008 (proj balance)334.01242413711.3942110.90.151.05
2008 (proj total)883.485857321421.21233723


Masa Kobayashi
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
20084203.0544.340281111.15182515
2008 (proj balance)2904.183029201011.281243141.130.131.26
2008 (proj total)7103.517469482121.2306829


Cliff Lee
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
200818182.31124.71091062001.03488532
2008 (proj balance)12123.668579582221.193318341.350.161.51
2008 (proj total)30302.852091881644221.18191366


Matt Ginter
YearGGSERAIPHitKBBHBPWHIPBFPHRERERA_deltaWHIP_deltaTotal_Delta
20081105550012000
2008 (proj balance)114.11332101.2814024.110.284.39
2008 (proj total)221.66887101.113402


Again, most Indians fans would agree that these are the pitchers who had good first halves. Rafael Perez, Aaron Laffey, and Fausto Carmona are all due to see their ERAs increase but their WHIPs decrease. I'm not sure what to make of that, other than the fact that neither is a perfect stat to indicate a pitcher's overall performance. Meanwhile, predictably, Cliff Lee is due for some regression to the mean. Still, I doubt anyone will be upset if he finishes with a 2.85 ERA and 1.10 WHIP.