Probability and statistics blog

Uncategorized — Comments Off
29
Jun 14

The week in stats (June 30th edition)

Normal Distribution vs. Paranormal Distribution

Professor Ramon van Handel of Princeton University posts his lecture notes on Probability in High Dimension.
Everday Anayltics shares his experience as a statistical consultant on a project for Delta Airline data using PCA and K-means Clustering.
If you know R and want to improve your SAS skills (or the other way around), check out this tutorial on logistic regressions using both packages.
Do you trade stocks? If so, here is A Simple Shiny App for Monitoring Trading Strategies.
Maybe I Don’t Really Know R After All.
And finally, why we should Separate Statistical Models of “What Is Learned” from “How It Is Learned”.

Uncategorized — Comments Off
22
Jun 14

The week in stats (June 2nd edition)

What people do with their degrees.

A sequence of 9 courses on Data Science will start on Coursera on 2 June and 7 July 2014, to be lectured by Professors of Johns Hopkins University. The courses are designed for students to learn to become Data Scientists and apply their skills in a capstone project. The courses are free, but if you want a Verified Certificate in the course, the Specialization Certificate or taking the Capstone Project, there is a small charge for that.
Do you know where people are going after college? Ben Schmidt, an assistant professor of history at Northeastern University, was curious about careers after college degrees, so he used a quick Sankey diagram to look at data from the American Community Survey.
Robert Seaton shares more than 100 interesting data sets for statistics. If you are looking for some numbers to get your hands dirty, this is the place to visit.
Cartesian Faith publishes the second chapter of his book (available for free on his website) called Modeling data with functional programming in R.
Nina Zumel from Win-Vector LLC discusses some useful tricks with the R function glm() in her recent blog post titled Trimming the Fat from glm() Models in R.
And finally, if you learned R in its early days (early 2000s or before), you may still be using some old-fashioned ways to accomplish some tasks better served by newer functions and packages. To help you become a better R coder, Revolution Analytics offers hipsteR: Learn what you missed in R as an early adopter.

Uncategorized — 1 comment
22
Jun 14

The week in stats (June 23rd edition)

If you are on the job market, Tal Galili from R bloggers has compiled 3 new R jobs for seekers like you.
Text mining is currently a live issue in data analysis. Enoromus text data resourses on the Internet made it an important component of Big Data world. If text mining is something that you need to do for your job, you should read Text mining in R – Automatic categorization of Wikipedia articles.
Randy Olsen, PhD student at Michigan State University’s Computer Science program, studies the percentages of undergraduate degrees conferred to men in the USA and publishes his findings in a blog titled The double-edged sword of gender equality.
Who will win the World Cup? See what statisticians say.
Earlier this month, the results of the 15th annual KDnuggets Software Poll were released and R’s popularity continues to grow. See Revolution Analytics’ new post for details.
And finally, Xi’an discusses a new paper by Simon Barthelmé and Nicolas Chopin called The Poisson transform for unnormalised statistical models.

Uncategorized — Comments Off
16
Jun 14

The week in stats (June 16th edition)

Writing functions is an important part of programming, and in order to write proper functions you need to know how to debug when your functions aren’t working. Slawa Rokicki, PhD student at Harvard, explains How to write and debug an R function.
It is often said that you should avoid loops in R because R is extremely slow with iterations, and hence many R-programmers try to avoid loops by working with matrices and arrays. Did you know that an even better option is to run your loops in C++ and import your result back into R? Here is a quick tutorial called how you can use C++ within R.
Rasmus Bååth blogs about the The Most Comprehensive Review of Comic Books Teaching Statistics.
Did you know that more and more startups are starting to use R as their primary data analysis tool? According to Revolution Analytics, Uber and CultureAmp have just joined the R camp.
Xi’an reviews a new paper called Generalizations related to hypothesis testing with the Posterior distribution of the Likelihood Ratio by Smith and Ferrari.
And finally, DiffusePrioR writes “If history can tell us anything about the World Cup, it’s that the host nation has an advantage of all other teams”. Do you agree or disagree, and what do you think is Brazil’s chance of winning the World Cup?

http://bit.ly/SHnlXH

Uncategorized — Comments Off
15
Jun 14

The week in stats (June 9th edition)

Like the plots above? Learn how to create these in R from Freakonometrics’ new post called Box plot, Fisher’s style.
If you are on the job market, Tal Galili from R bloggers has compiled 6 new R jobs for seekers like you.
Big Data has gained lots of popularity recently, and every data scientist should know at least something about it. If you are new to data science, consider this introduction to R for Big Data with PivotalR.
Using Repeated Measures to Remove Artifacts from Longitudinal Data by Dmitry Grapov.
And finally, Andrew Gelman discusses Why we hate stepwise regression.

art / epistomology / stats — 2 comments
12
Jun 14

A new way to visualize content

Right now I’m working on a project that involves new ways to view units of content and the relationships between them. I’ve posted the comic I worked on, it has a number of stats references throughout. This is early alpha stages for the software, you may run into issues. To see the relationships, go to the puffball menu and make sure that “Show relationships” is clicked.

Uncategorized — Comments Off
26
May 14

The week in stats (May 26th edition)

Alvaro Galindo reviews Social Media Mining with R by by Nathan Danneman and Richard Heinmann.
Some popular articles on R tip and tricks are: R has some sharp corners by Win-Vector LLC, Sample uniformly within a fixed radius by Forester (Assistant Professor at the University of Minnesota Twin Cities), The Birthday Simulation by Wes Stevenson, and didYouMean() Function: Using Google to correct errors in Strings by Sam Weiss.
R bloggers compiles a list of R related positions for those who are on the job market.
Xi’an discusses a special issue Statistical Science named Big Bayes Stories: A Collection of Vignettes.
Last week, we featured an article on R vs. Julia. This week, Matloff (aka Mad (Data) Scientist) writes another comparison called R beats Python! R beats Julia! Anyone else wanna challenge R?

Uncategorized — Comments Off
19
May 14

The week in stats (May 19th edition)

Which Team Do You Cheer For? An N.B.A. Fan Map by The New York Times

Are you a self-taught “scientist programmer”? Here is why people think code written by people like you is ugly.
As always, R articles are extremely popular. This week, we have: Facebook teaches you exploratory data analysis with R by Revolution Analytics, Beyond R, or on the Hunt for New Tools by Quintuitive, Bootstrap Critisim (with example) by Eran Raviv, The apply command 101 by Learning R by Imitation, and Can We do Better than R-squared? by Learning as You Go.
Julia is a new programming language (only 2 years old) for scientific computing and it has gained lots of popularity recently. In the past, we shared some articles comparing R and Julia. This week, Alvaro Galindo writes another comparison called Julia versus R – Playing around.
Sébastien Bubeck, assistant professor at Princeton, releases the first draft of his monograph based some old lecture notes called Theory of Convex Optimization for Machine Learning.
And finally, happy Victoria Day to those in Canada!

Uncategorized — 2 comments
12
May 14

The week in stats (May 12th edition)

Type I and Type II errors explained.

Looking for a job? Here are some jobs compiled by R-bloggers that may be of interest to you.
Homer White, professor of mathematics at Georgetown College, shares his Five Reasons to Teach Elementary Statistics With R.
Seven R Quirks That Will Drive You Nutty.
Some popular R articles this week are: how to build a sales dashboard with R, Optimising your R code, and Modelling seasonal data with GAMs.
And finally, Xi’an discusses bridging the gap between machine learning and statistics.

Uncategorized — Comments Off
05
May 14

The week in stats (May 5th edition)

Popular R articles this week are: colormap by Dan Kelley (Professor of Oceanography at Dalhousie University), The new look of learning R by DataCamp, Writing an R package from scratch by Hilary Parker (Data Analyst at Etsy), Test coverage of the 10 most downloaded R packages by Quartz Bio, How to Code Something ‘New’ in R by Francis Smart (PhD student at Michigan State University) and Reading large data tables in R by Fabio Marroni.
If you roll a fair die 6 times, what is the probability that there is at least one pair of identical consecutive face values?
Great news! The RSS is setting a data analysis challenge this year. The top three teams will be invited to present their results in a special session at the RSS Annual Conference in September 2014, and submissions will be considered for publication in the Journal of the Royal Statistical Society, Series C. If you are interested, here are the details.
And finally, do you fly frequently? If so, you may want to know how to Automatically Scrape Flight Ticket Data Using R and Phantomjs.

Probability and statistics blog

The week in stats (June 30th edition)

The week in stats (June 2nd edition)

The week in stats (June 23rd edition)

The week in stats (June 16th edition)

The week in stats (June 9th edition)

A new way to visualize content

The week in stats (May 26th edition)

The week in stats (May 19th edition)

The week in stats (May 12th edition)

The week in stats (May 5th edition)

Recent Posts

Recent Comments

Archives

Categories

Meta