Maybe I Don’t Really Know R After All

nested_r_dataframe

Lately, I've been feeling that I'm spreading myself too thin in terms of programming languages. At work, I spend most of my time in Hive/SQL, with the occasional Python for my smaller data. I really prefer Julia, but I'm alone at work on that one. … [Continue reading]

Using Julia As A ‘Glue’ Language

While much of the focus in the Julia community has been on the performance aspects of Julia relative to other scientific computing languages, Julia is also perfectly suited to 'glue' together multiple data sources/languages. In this blog post, I will … [Continue reading]

Five Hard-Won Lessons Using Hive

I've been spending a ton of time lately on the data engineering side of 'data science', so I've been writing a lot of Hive queries. Hive is a great tool for querying large amounts of data, without having to know very much about the underpinnings of … [Continue reading]