DiscoverData Skeptic
Data Skeptic
Claim Ownership

Data Skeptic

Author: Kyle Polich

Subscribed: 30,205Played: 445,922


The Data Skeptic Podcast features interviews and discussion of topics related to data science, statistics, machine learning, artificial intelligence and the like, all from the perspective of applying critical thinking and the scientific method to evaluate the veracity of claims and efficacy of approaches.
369 Episodes
Shereen Elsayed and Daniela Thyssens, both are PhD Student at Hildesheim University in Germany, come on today to talk about the work “Do We Really Need Deep Learning Models for Time Series Forecasting?”
Detecting Drift

Detecting Drift


Sam Ackerman, Research Data Scientist at IBM Research Labs in Haifa, Israel, joins us today to talk about his work Detection of Data Drift and Outliers Affecting Machine Learning Model Performance Over Time.
Julien Herzen, PhD graduate from EPFL in Switzerland, comes on today to talk about his work with Unit 8 and the development of the Python Library: Darts. 
Welcome to Timeseries! Today’s episode is an interview with Rob Hyndman, Professor of Statistics at Monash University in Australia, and author of Forecasting: Principles and Practices.
Today's experimental episode uses sound to describe some basic ideas from time series. This episode includes lag, seasonality, trend, noise, heteroskedasticity, decomposition, smoothing, feature engineering, and deep learning.  
Orders of Magnitude

Orders of Magnitude


Today’s show in two parts. First, Linhda joins us to review the episodes from Data Skeptic: Pilot Season and give her feedback on each of the topics. Second, we introduce our new segment “Orders of Magnitude”. It’s a statistical game show in which participants must identify the true statistic hidden in a list of statistics which are off by at least an order of magnitude. Claudia and Vanessa join as our first contestants.  Below are the sources of our questions. Heights Bird Statistics Birds in the US since 2000 Causes of Bird Mortality Amounts of Data Our statistics come from this post
AI has, is, and will continue to facilitate the automation of work done by humans. Sometimes this may be an entire role. Other times it may automate a particular part of their role, scaling their effectiveness. Unless progress in AI inexplicably halts, the tasks done by humans vs. machines will continue to evolve. Today’s episode is a speculative conversation about what the future may hold. Co-Host of Squaring the Strange Podcast, Caricature Artist, and an Academic Editor, Celestia Ward joins us today! Kyle and Celestia discuss whether or not her jobs as a caricature artist or as an academic editor are under threat from AI automation. Mentions The legendary Dr. Jorge Pérez and his work studying unicorns Supernormal stimulus International Society of Caricature Artists Two Heads Studios
Today on the show Derek Driggs, a PhD Student at the University of Cambridge. He comes on to discuss the work Common Pitfalls and Recommendations for Using Machine Learning to Detect and Prognosticate for COVID-19 Using Chest Radiographs and CT Scans. Help us vote for the next theme of Data Skeptic! Vote here:
Given a document in English, how can you estimate the ease with which someone will find they can read it?  Does it require a college-level of reading comprehension or is it something a much younger student could read and understand? While these questions are useful to ask, they don't admit a simple answer.  One option is to use one of the (essentially identical) two Flesch Kincaid Readability Tests.  These are simple calculations which provide you with a rough estimate of the reading ease. In this episode, Kyle shares his thoughts on this tool and when it could be appropriate to use as part of your feature engineering pipeline towards a machine learning objective. For empirical validation of these metrics, the plot below compares English language Wikipedia pages with "Simple English" Wikipedia pages.  The analysis Kyle describes in this episode yields the intuitively pleasing histogram below.  It summarizes the distribution of Flesch reading ease scores for 1000 pages examined from both Wikipedias.  
Today on the show we have Shubhranshu Shekar, a Ph. D Student at Carnegie Mellon University, who joins us to talk about his work, FAIROD: Fairness-aware Outlier Detection.
Life May be Rare

Life May be Rare


Today on the show Dr. Anders Sandburg, Senior Research Fellow at the Future of Humanity Institute at Oxford University, comes on to share his work “The Timing of Evolutionary Transitions Suggest Intelligent Life is Rare.” Works Mentioned: Paper: “The Timing of Evolutionary Transitions Suggest Intelligent Life is Rare.”by Andrew E Snyder-Beattie, Anders Sandberg, K Eric Drexler, Michael B Bonsall  Twitter: @anderssandburg
Social Networks

Social Networks


Mayank Kejriwal, Research Professor at the University of Southern California and Researcher at the Information Sciences Institute, joins us today to discuss his work and his new book Knowledge, Graphs, Fundamentals, Techniques and Applications by Mayank Kejriwal, Craig A. Knoblock, and Pedro Szekley. Works Mentioned “Knowledge, Graphs, Fundamentals, Techniques and Applications”by Mayank Kejriwal, Craig A. Knoblock, and Pedro Szekley
The QAnon Conspiracy

The QAnon Conspiracy


QAnon is a conspiracy theory born in the underbelly of the internet.  While easy to disprove, these cryptic ideas captured the minds of many people and (in part) paved the way to the 2021 storming of the US Capital. This is a contemporary conspiracy which came into existence and grew in a very digital way.  This makes it possible for researchers to study this phenomenon in a way not accessible in previous conspiracy theories of similar popularity. This episode is not so much a debunking of this debunked theory, but rather an exploration of the metadata and origins of this conspiracy. This episode is also the first in our 2021 Pilot Season in which we are going to test out a few formats for Data Skeptic to see what our next season should be.  This is the first installment.  In a few weeks, we're going to ask everyone to vote for their favorite theme for our next season.  
Karthick Shankar, Masters Student at Carnegie Mellon University, and Somali Chaterji, Assistant Professor at Purdue University, join us today to discuss the paper "JANUS: Benchmarking Commercial and Open-Source Cloud and Edge Platforms for Object and Anomaly Detection Workloads" Works Mentioned: “JANUS: Benchmarking Commercial and Open-Source Cloud and Edge Platforms for Object and Anomaly Detection Workloads.” by: Karthick Shankar, Pengcheng Wang, Ran Xu, Ashraf Mahgoub, Somali ChaterjiSocial Media Karthick Shankar Somali Chaterji
Hal Ashton, a PhD student from the University College of London, joins us today to discuss a recent work Causal Campbell-Goodhart’s law and Reinforcement Learning. "Only buy honey from a local producer." - Hal Ashton   Works Mentioned: “Causal Campbell-Goodhart’s law and Reinforcement Learning”by Hal AshtonBook  “The Book of Why”by Judea PearlPaper Thanks to our sponsor!  When your business is ready to make that next hire, find the right person with LinkedIn Jobs. Just visit to post a job for free! Terms and conditions apply
Yuqi Ouyang, in his second year of PhD study at the University of Warwick in England, joins us today to discuss his work “Video Anomaly Detection by Estimating Likelihood of Representations.”Works Mentioned: Video Anomaly Detection by Estimating Likelihood of Representations by: Yuqi Ouyang, Victor Sanchez
Nirupam Gupta, a Computer Science Post Doctoral Researcher at EDFL University in Switzerland, joins us today to discuss his work “Byzantine Fault-Tolerance in Peer-to-Peer Distributed Gradient-Descent.”   Works Mentioned: Byzantine Fault-Tolerance in Peer-to-Peer Distributed Gradient-Descent by Nirupam Gupta and Nitin H. Vaidya   Conference Details:
Mikko Lauri, Post Doctoral researcher at the University of Hamburg, Germany, comes on the show today to discuss the work Information Gathering in Decentralized POMDPs by Policy Graph Improvements. Follow Mikko: @mikko_lauri Github
Leaderless Consensus

Leaderless Consensus


Balaji Arun, a PhD Student in the Systems of Software Research Group at Virginia Tech, joins us today to discuss his research of distributed systems through the paper “Taming the Contention in Consensus-based Distributed Systems.”  Works Mentioned “Taming the Contention in Consensus-based Distributed Systems”  by Balaji Arun, Sebastiano Peluso, Roberto Palmieri, Giuliano Losa, and Binoy Ravindran “Fast Paxos” by Leslie Lamport
Maartje ter Hoeve, PhD Student at the University of Amsterdam, joins us today to discuss her research in automated summarization through the paper “What Makes a Good Summary? Reconsidering the Focus of Automatic Summarization.”  Works Mentioned  “What Makes a Good Summary? Reconsidering the Focus of Automatic Summarization.” by Maartje der Hoeve, Juilia Kiseleva, and Maarten de Rijke Contact Email: Twitter: Website:
Comments (18)



Jan 23rd


@6:00: The threshold for statistical significance does not "depend on the outcome." It raises a red flag even to hear someone say that, especially the host of a "data science" podcast. (Of course, if he knew what he was talking about, he'd be a "statistician" instead.) He might more accurately have said that any such estimate of the minimum sample size depends on the number of planned comparisons and the assumed effect size for each measured effect. Confusion about this should disqualify someone from hosting such a podcast.

Aug 23rd


@2:19: Too much interpretation as if respondents were randomly sampled. Respondents self-selected.

Aug 23rd

Antonio Andrade

thanks so much for sharing the results

Aug 12th


@1:03: It doesn't "beg the question"; it "raises the question." To "beg the question" is to commit a logical fallacy in which one assumes the conclusion.

Jun 15th

Benjamin Weckerle

Is the spin-off / journal club podcast on castbox?

Jun 2nd

Platte Gruber

KILLER intro, awesome work!

Jan 8th
Reply (1)

Marco Gorelli

"I find it stunning that people don't do that. The only thing I can think of is that there's just the lack of focused time. There's so many things we could spend our time on now we spend a little on all of them and we don't have depth that we need. A lot of people will come to a conference or something like that just to be away from work and only focus on one thing. Unfortunately they also bring their phone and completely break that paradigm. " ouch

Dec 8th

Bhavul Gauri

Brilliantly put!

Aug 27th

Akshay Shirsath

Thoughtful episode.

Jun 8th
Reply (1)

Achint Verma

A very very high level introduction to Kalman Filters. You could have talked about the matrices.

Mar 29th

Troy Kirin

Golden, thanks for this!

Mar 19th

Vannucci Santos

Why the guy talking about ethics was so evasive?

Dec 25th

Anna Malahova

I love everything about this podcast channel! It is easy to listen to, easy to understand without data science background, interesting topics and examples of situations where to apply. Very enjoyable and entertaining delivery. Informative show notes that help you to recall what the episode was about even after a while. Really can't think about any downsides. I am listening to all episodes starting from early days like an audiobook, love how music into evolved over time.

Nov 24th

Abdul Wahab Abrar

What about using Deep Learning techniques directly and integrate it with Neuroscience

Feb 12th

Giancarlo Vercellino

rilevante dal 23esimo minuto

Dec 12th
Download from Google Play
Download from App Store