PinnedAdir Mashiach·Jul 31, 2019Apache Spark: 5 Performance Optimization TipsInteresting and specific lessons learned from experience
PinnedAdir Mashiach·May 27, 2020Partition Management in HadoopOur solution to the Hadoop small files problem
PinnedAdir Mashiach·Feb 14, 2019Defend Your Infrastructure — Handling 3,000 Hungry UsersWhy is it so important to track your users’ queries, and how we do it?
Adir Mashiach·May 23, 2021Firebolt — The new kid on the (data warehousing) blockA short description of Firebolt, for not-so-technical people
Adir Mashiach·Dec 26, 2020SPOT: Is Spotify a good stock to buy?A human-readable stock analysis, from a rational perspective
Adir Mashiach·Aug 31, 2018Impala Discussion With The Product Manager (Greg Rahn)Q&A session on specific issues that bothered usA response icon1A response icon1
Adir Mashiach·Aug 15, 20185 Main Missing Features in Impala (Opinion)A letter to the developers and product manager of ImpalaA response icon3A response icon3
Adir Mashiach·Apr 13, 2018Hotspotting In Hadoop — Impala Case StudyWhy Small Frequently-Queried Tables Shouldn’t Be Stored In HDFS?A response icon2A response icon2
Adir Mashiach·Apr 1, 2018Partition Index - Selective Queries On Really Big TablesHow to make your selective queries run 100x faster?A response icon1A response icon1
Adir Mashiach·Mar 20, 2018Apache Impala: My Insights and Best PracticesHow did we make our Impala run faster?A response icon3A response icon3