r/learndatascience 5d ago

Resources I made a structured AI/ML playlist that moves from foundations toward deployment

2 Upvotes

I created a free YouTube learning playlist for people who want to see the full AI/ML journey rather than only isolated model demos. The series is organized from foundations through data preparation, evaluation, practical modeling, and deployment-oriented topics, with lessons intended to be followed in sequence.

Playlist: https://www.youtube.com/playlist?list=PLUccxCq754BE

I’m sharing it transparently because I run the channel. I’d value feedback from learners: where do most beginner courses lose you — data preparation, model evaluation, or deployment? If this kind of resource is not appropriate as a standalone post, I’ll remove it and use the community’s preferred resource thread.


r/learndatascience 6d ago

Personal Experience Always 2 full time jobs or 1 prestigious

2 Upvotes

Guys so i am fresh from uni and observe two kind of tracks of my peers in data science. They go with the two chill full time jobs (1 is remote) and get promotions at both work places but both are chill ones, or they go in 1 company which is considered prestigious and overwork till 10-11 pm and get promotions there, from your experience which kind of way leads to a better quality of life e.g. work life balance wise, career wise, experience wise and income wise? Should i find 1 prestigious place and hope for the long run and work there or work at 2 chill places. Both seem to obtain promotions in similar periods but the ones with two jobs seem to be more happy and even get more money. I decided to pursue masters and only now 1 year after them decided to join workforce but cannot distinguish which way is better, could you give your advice


r/learndatascience 6d ago

Resources I’m writing an AI Safety book for people who actually want to do the technical work

4 Upvotes

I’ve been reading quite a lot around AI safety, and one thing I’ve found frustrating is how difficult it is to find material that bridges the gap between talking about AI safety and actually doing AI safety research.

There are plenty of good books and papers discussing the political, socioeconomic, philosophical and governance questions around AI. Those conversations are obviously important but if you’re a data scientist, ML engineer or someone already working with AI few tell you how doing AI safety research actually looks like in Python?

For example how to actually run an experiment looking at sycophancy, how to create or use a dataset to test whether a model changes its answer because of information about the user, how to measure the behaviour, what does the evaluation code look like, how do you interpret the results, and what can you not conclude from them?

That is the gap I’m trying to address with a book I’m currently writing.

The approach is very practical. Each topic starts with the safety problem and the research behind it, but then we actually build the experiment. Python code, datasets, models, metrics, results and discussion of the limitations, all those.

The idea isn’t to pretend that running a few notebooks suddenly makes someone an AI safety researcher. It’s to make the field much more approachable to people who already have data science or AI skills and want to understand what technical AI safety research actually involves.

I wrote a technical book on Practical LLMOps and this was one of the chapters. Because of the sheer amount of content, I stripped it bare in that book and made it its own.

I’ve just published Chapter 1 on Substack (just finished chapter 4). It introduces the approach I’m taking with the book and starts building that bridge between AI safety as a subject people discuss and AI safety as something we can actually investigate experimentally.

Would genuinely be interested in feedback, particularly from people already working in ML, data science or AI safety.

https://open.substack.com/pub/houstonmuzamhindo/p/i-am-writing-a-practical-technical


r/learndatascience 5d ago

Question Is a 50k PKR Data Science Course Worth Going Into Debt For?

Thumbnail
1 Upvotes

r/learndatascience 6d ago

Question Programing problem

Thumbnail
1 Upvotes

I am learning in data scientist. I have learned python basics, pandas, numpy, etc , sql ,ml dl, nlp but i struggle in problem solving what should i do


r/learndatascience 6d ago

Question Just how far behind am I, or where exactly do I stand right now?

12 Upvotes

I I’ve learned Python, SQL, Excel, NumPy, Pandas, Matplotlib, Seaborn, and basic Git/GitHub. I’m now moving into Machine Learning.

Am I behind, on track, or ahead compared with other 2nd-year students?

What skills should someone seriously pursuing Data Science/ML have by the end of their second year?

I’d appreciate honest feedback, especially from people working or interning in the field.


r/learndatascience 6d ago

Career Best data science course

6 Upvotes

I want to pursue data science course can anybody which one is best institute for data science course
Please everybody give suggestion as i am new to it


r/learndatascience 6d ago

Question Decided to dive in ds

1 Upvotes

So guys hi, just worry a lot, decided to pursue career in dd after bachelor in economics, now kinda studied a big chubk of ds already but i am astounded at the amount of stuff that you should literally remember. Any tips on how to memorize all the details of every neural network and characteristics of distributions? Also i suck at coding but since I will have to work on deployment and services as well, i understand that i have to push that field too. I found a yandex course to be of a little use since the problems described there are isolated and does not help to get overall picture of how to code in ds. Any tips on how to reinforce all this knowledge? Cause i am at a loss right now


r/learndatascience 6d ago

Question Recource Confusion, Self Learning paced, Progression, and Community!

Thumbnail
1 Upvotes

r/learndatascience 6d ago

Question I want to start my Data Analysis I want the perfect crash course

0 Upvotes

I want to start learning data analysis and i have some knowledge on data science and ML , which free resources do you recommend and crash course would be better.


r/learndatascience 6d ago

Resources I've built a library of concept notes on over 150+ DS topics - with a focus on revision and retention

1 Upvotes

I've worked in data science for years and I still forget things constantly. Courses never fixed that — I'd finish one, feel like I'd learned it, then google the same concept a month later when it actually came up at work. NotebookLM's way of structuring material around you was the closest thing I found, but I still had to go and find the sources myself.

So I built the thing I wanted. DS concepts broken into small nodes arranged as a mind map rather than a linear syllabus, each with code examples and a practice section — about 150 topics so far. The part I care most about: once you complete a topic it starts decaying on a forgetting-curve schedule, and the map visibly goes cold. When it does, you get a review slice — flash cards and a short quiz — targeted at what you've actually lost rather than what's next in a queue.

Some topics to start with -

  1. https://www.bitelrn.com/library/linear-algebra-matrices

  2. https://www.bitelrn.com/library/principal-component-analysis

  3. https://www.bitelrn.com/library/sql-basics

Full app: https://www.bitelrn.com — the first phase is permanently free including the decay and review mechanics; later phases are paid. Saying that upfront so nobody feels ambushed.

I'm curious what actually works for other people here. Do you make notes, bookmark links, or just re-google it every time?


r/learndatascience 7d ago

Question Breaking into Data

8 Upvotes

A couple of weeks ago I asked a sub what would be the best way of breaking in data and what would be worth looking into. I got some pretty decent feedback and decided to segue into something a bit different

Being that I own AI licenses, I have spent the last 2-3 weeks developing a full stack data engineering course from beginner to advanced. Now I wanted to know if it was possible to ask if anyone would like to take a look into it, give me some feedback and recommendations, and an overall rating.

I can handle criticism don’t worry, I just want to develop something, learn the info from what I’ve built and then eventually publish it publicly for people to have a free resource to utilize.

(This is not promo, rather I want to gain feedback from actual analysts and engineers to see if this is a good platform)


r/learndatascience 7d ago

Question [D] Looking for advice: Modelling a medicine-reminder agent that must decide “remind / wait / notify” under incomplete information[D]

Thumbnail
1 Upvotes

r/learndatascience 8d ago

Question How should feature selection be handled across different forecasting and predictive models?

4 Upvotes

I’m working on a forecasting / predictive modeling problem where I’m comparing four groups of models:

Baseline: Naive / Seasonal Naive

Time Series: ARIMA, SARIMA, ETS, Prophet, etc.

Statistical: Linear Regression, Ridge, GAM, etc.

Machine Learning: Random Forest, Gradient Boosting, XGBoost, etc.

One issue I’m struggling with is feature selection.

Should I create one universal feature subset before modeling and give the same predictors to the Statistical and ML models, or should feature selection be model-specific?

For example, a linear model may benefit from correlation filtering, VIF, or LASSO, while tree-based models may select a different set of variables because they can capture nonlinearities and interactions.

I’m also unsure about the correct modeling pipeline. Where exactly should feature selection happen relative to the train/test split, cross-validation, preprocessing, and hyperparameter tuning?

Some questions I’m trying to resolve:

Should all comparable models use the same feature set for a fair model comparison?

If model-specific feature selection is appropriate, how should it be implemented without introducing data leakage?

Should feature selection be performed separately within each model’s training process?

If different models select different predictors, how should that variation be explained to stakeholders?

Which feature selection or interpretation methods are actually useful for identifying the underlying drivers of the dependent variable, rather than simply variables that improve predictive accuracy?

How should I distinguish between predictive importance and explanatory importance when interpreting the results?

Ultimately, I’m trying to build a workflow that balances forecasting accuracy, fair model comparison, and interpretability.

What would be a statistically sound end-to-end approach for handling feature selection across these different model types?


r/learndatascience 8d ago

Question Where can I find a good dataset for a real-world data science project?

22 Upvotes

I’m trying to build a proper data science/ML project, but I’m having a hard time finding a dataset that is big enough and not already used by everyone.
For example, there are datasets like the UK Online Retail dataset, Olist, and other popular sales/retail datasets. They’re good datasets, but I see them being used in a lot of projects already.
I don’t want to just download a dataset, do some EDA, train a model and put it on my resume. I want to build something around an actual business problem, where I have to figure out what the problem is, analyze the data, come up with useful insights, maybe build a model, and actually explain how it could help the business.
So where do you guys usually find datasets for this?
Should I try to find data from smaller companies, government sources, APIs, research papers, etc.? Or is it okay to create my own dataset using AI/cloud tools and then create a realistic business problem around it?
For example, if I create a large synthetic sales dataset, could I create a realistic business scenario around it and then treat it like a real project — forecasting sales, understanding customer behavior, optimizing inventory, etc.?
Would that be considered a decent portfolio project, or is using real-world data much better?
I’d mainly like to hear from people who have built projects for their portfolios or have experience hiring for data science/ML roles. Where do you actually get your data from when you want to build something that’s not the same Kaggle project everyone has already done?


r/learndatascience 8d ago

Resources Beginner friendly AI & ML Videos

Thumbnail
youtube.com
5 Upvotes

When I was a student, I often needed very simple machine learning explanations before exams not a full course, not heavy math from the first minute, just someone explaining the intuition clearly.

That’s why I started making short beginner-friendly ML videos.

The idea is to explain topics in a simple visual way first.

I’m not trying to replace proper courses or textbooks. I’m trying to make the “okay, what is actually happening here?” part easier to understand.


r/learndatascience 8d ago

Original Content Running a free session on what AI agents actually do to a real data science workflow, happy to share details in comments

Thumbnail
1 Upvotes

r/learndatascience 9d ago

Question How do you decide when a machine learning model is good enough?

3 Upvotes

When building a machine learning model, how do you decide that the results are good enough to move forward?

Do you mainly look at metrics like accuracy, precision/recall, F1, or RMSE, or do you also consider things like the business problem, model complexity, and how well it performs on unseen data?

I'm curious how people approach this when working on real projects rather than just following a tutorial.


r/learndatascience 9d ago

Resources Python Dictionary Mutation and Copying

Post image
6 Upvotes

An exercise to help build the right mental model for Python data. - Solution - Explanation - More exercises

The “Solution” link uses package memory_graph to visualize execution and reveals what’s actually happening.


r/learndatascience 9d ago

Discussion Kinda stuck at EDA in ML

Thumbnail
1 Upvotes

r/learndatascience 9d ago

Question How do you get better at the practical side of Data Science? I feel stuck when it comes to coding

10 Upvotes

Hi everyone,

I am currently preparing to switch into Data science. The theory side is mostly fine for me but I am really struggling to apply what I learn practically, especially when coding.

I have been working with datasets in kaggle and trying to practice but I often feel stuck in :

what should I do first

what should be my next step

What code should I implement

What questions should I ask the data

EDA is where I struggle the most and I often jump to ChatGPT and asking what should i do next.

I dont want to be the person who just follow tutorials and copy paste the code. I want to develop the ability to look at a problem and think of a solution myself and write code without any help.

I’d really appreciate any advice. Thank you!


r/learndatascience 9d ago

Question this Old model laptop efficient for a data science student?

Post image
2 Upvotes

r/learndatascience 10d ago

Question Machine Learning roadmap & guidance

Thumbnail
1 Upvotes

r/learndatascience 11d ago

Question Looking for Current Students/Alumni – MSc Statistics (Co-op), University of Regina

1 Upvotes

Looking for Current Students/Alumni – MSc Statistics (Co-op), University of Regina

Hi everyone! I’ll be starting my MSc in Statistics (Co-op) at the University of Regina this Fall 2026 and am planning to arrive in Regina in the first week of September.

I’d really appreciate hearing from any current students or alumni of the MSc Statistics program, especially those who are starting/started the program recently.

I’d love to know:

What are the courses like, especially the workload and difficulty?

How is the overall experience in the Statistics department?

What should I expect from the Co-op component?

Any advice for preparing before classes start?

If anyone else is joining the Statistics program this Fall, feel free to connect as well. It would be great to know some fellow students before arriving!

Thanks in advance! 🙏


r/learndatascience 11d ago

Question Which version of Python do I need to use for Data Science?

2 Upvotes

Should I use the 3.14 version?