this post was submitted on 01 Dec 2024

37 points (100.0% liked)

Programming

427 readers

1 users here now

Welcome to the main community in programming.dev! Feel free to post anything relating to programming here!

Cross posting is strongly encouraged in the instance. If you feel your post or another person's post makes sense in another community cross post into it.

Hope you enjoy the instance!

Rules

Follow the programming.dev instance rules
Keep content related to programming in some way
If you're posting long videos try to add in some form of tldr for those who don't want to watch videos

Wormhole

Follow the wormhole through a path of communities !webdev@programming.dev

founded 2 years ago

MODERATORS

snowe@programming.dev

Ategon@programming.dev

MaungaHikoi@lemmy.nz

Any data scientists out there? What's your go to programming language and tools for your work? (lm.paradisus.day)

submitted 4 months ago by rutrum@lm.paradisus.day to c/programming@programming.dev

15 comments fedilink hide all child comments

No surprise I use python, but I've recently started experimenting with polars instead of pandas. I've enjoyed it so far, but Im not sure if the benefits for my team's work will be enough to outweigh the cost of moving from our existing pandas/numpy code over to polars.

I've also started playing with grafana, as a quick dashboarding utility to make some basic visualizations on some live production databases.

top 15 comments

sorted by: hot top controversial new old

[–] ptz@dubvee.org 10 points 4 months ago

I'm not a data scientist but I support a handful. They all use Python for the most part, but a few of them (still?) use R. Then there's the small group that just throws everything into Excel 🤷🏻‍♂️

[–] Felling_High_Horses@endlesstalk.org 5 points 4 months ago

R is my go-to, since that's what my uni taught me (Utrecht university). But I've been learning pandas on python on the side for the versatility (and my CV).

[–] driving_crooner@lemmy.eco.br 4 points 4 months ago (1 children)

Not a data scientist, but an actuarie. I use python, pandas in jupyter notebooks (vs code). I think it would be cool to use polars, but my datasets are not that big to justify the move.

[–] rutrum@lm.paradisus.day 1 points 4 months ago

If it works, don't fix it!

[–] magic_lobster_party@fedia.io 4 points 4 months ago

Java with Spark.

Although I feel like I’m doing less of data science and more of data processing.

[–] SplashJackson@lemmy.ca 3 points 4 months ago (1 children)

I like pandas but sometimes figuring out the simplest of shit is so complicated

[–] rutrum@lm.paradisus.day 1 points 4 months ago

I learned SQL before pandas. It's still tabular data, but the mechanisms to mutate/modify/filter the data are different methodologies. It took a long time to get comfy with pandas. It wasnt until I understood that the way you interact with a database table and a dataframe are very different, that I started to finally get a grasp on pandas.

[–] verdeviento@mander.xyz 2 points 4 months ago (1 children)

What do you enjoy/find beneficial about polars?

[–] rutrum@lm.paradisus.day 3 points 4 months ago (1 children)

Its a paradigm shift from pandas. In polars, you define a pipeline, or a set of instructions, to perform on a dataframe, and only execute them all at once at the end of your transformation. In other words, its lazy. Pandas is eager, which every part of the transformation happens sequentially and in isolation. Polars also has an eager API, but you likely want to use the lazy API in a production script.

Because its lazy, Polars performs query optimization, like a database does with a SQL query. At the end of the day, if you're using polars for data engineering or in a pipeline, it'll likely work much faster and more memory efficient. Polars also executes operations in parallel, as well.

[–] Kache@lemm.ee 1 points 4 months ago (1 children)

What kind of query optimization can it for scanning data that's already in memory?

[–] rutrum@lm.paradisus.day 5 points 4 months ago (1 children)

A big feature of polars is only loading applicable data from disk. But during exporatory data analysis (EDA) you often have the whole dataset in memory. In this case, filters wont help much there. Polars has a good page in their docs about all the possible optimizations it is capable of. https://docs.pola.rs/user-guide/lazy/optimizations/

One I see off the top is projection pushdown, which only selects relevant columns for a final transformations. In pandas, if you perform a group by with aggregation, then only look at a few columns, you still perform aggregation across all the data. In polars lazy API, you would define the entire process upfront, and it would know not to aggregate certain columns, for instance.

[–] Kache@lemm.ee 1 points 4 months ago* (last edited 4 months ago) (1 children)

Hm, that's kind of interesting

But my first reaction is that optimizations only at the "Python processing level" are going to be pretty limited since it's not going to have metadata/statistics, and it'd depend heavily on the source data layout, e.g. CSV vs parquet

[–] rutrum@lm.paradisus.day 1 points 4 months ago

You are correct. For some data sources like parquet it includes some metadata that helps with this, but it's not as robust at databases I dont think. And of course, cvs have no metadata (I guess a header row.)

The actually specification for how to efficiently store tabular data in memory that also permits quick execution of filtering, pivoting, i.e. all the transformations you need...is called apache arrow. It is the backend of polars and is also a non-default backend of pandas. The complexity of the format I'm unfamiliar with.

[–] milicent_bystandr@lemm.ee 2 points 4 months ago

I only dabble, but I really like Julia. Has several language and architecture features I really like compared to python. Also looks like the libraries have been getting really good since last I used it much.

[–] odium@programming.dev 1 points 4 months ago

data engineer, not scientist here. Mostly Python and pyspark for me.