Hacking Academia: Data Science and the University
A reflection on the challenges of data science in academia, discussing the 'brain drain' of data skills and the need for systemic change.
A reflection on the challenges of data science in academia, discussing the 'brain drain' of data skills and the need for systemic change.
A report on the 2014 scikit-learn developer sprint in Paris, covering participants, venues, achievements, and sponsors.
Explores how personas, data science, and k-means clustering can be used together to analyze user data and gain actionable business insights.
A guide to using the Unix command-line for efficient data science workflows, including data processing, exploration, and modeling.
A guide to setting up a remote IPython Notebook server on Amazon EC2 for data science and analytics.
Explores how the demand for big data skills in industry is draining talent from academic science, threatening research.
A guide to seven essential command-line tools (jq, csvkit, Rio, etc.) for data scientists to obtain, scrub, explore, and model data.
Argues for the importance of statistical theory in data science, using examples from medical research to show where abstract theory solved practical problems.
A data scientist analyzes gender stereotypes in children's clothing by building a model to classify t-shirts using image processing and color data.
Argues that true data science and innovation require deep mathematical understanding, not just push-button tools, and defends the value of skilled data scientists.
A review of the book 'NumPy 1.5 Beginner's Guide', covering its content, style, and suitability for learning numerical computing with Python.
A recap of EuroSciPy 2010, highlighting the growth of the Python in science conference, key topics, and the community atmosphere.
Announcing EuroScipy 2010, a European conference on Python for scientific computing, featuring tutorials and keynote speakers in Paris.