Databricks

Spark’s Coalesce vs Repartition vs Repartition-by-Range – My Experience with them

Spark’s Coalesce vs Repartition vs Repartition-by-Range – My Experience with them   If you’ve spent any time tuning Spark jobs, you’ve run into the classic question: do I call `coalesce()`, `repartition()`, or `repartitionByRange()`? All three change how your data is partitioned across the cluster, but they behave very differently under the hood — and choosing

Spark’s Coalesce vs Repartition vs Repartition-by-Range – My Experience with them Read More »

From Custom Docker Images to One‑Click Libraries: My Experience Customizing AWS EMR Serverless vs Azure Databricks Compute

Modern data platforms live or die by how quickly you can ship code to production.For one of our recent projects, that speed was determined by something deceptively simple: adding a custom Python module to our distributed jobs. We started on AWS EMR Serverless and later moved to Azure Databricks for compute jobs.Both platforms can absolutely

From Custom Docker Images to One‑Click Libraries: My Experience Customizing AWS EMR Serverless vs Azure Databricks Compute Read More »