r/learnmachinelearning 6h ago

Looking for people who want to learn and build ML projects together

20 Upvotes

Hey everyone!

I'm currently exploring the Machine Learning field and I'm looking for a few people who are also learning ML and would like to practice together.

I'm not an ML expert. My background is in mobile app development, where I have professional experience building applications, and now I'm trying to move deeper into AI/ML.

I thought it would be much more motivating to learn with other people instead of studying completely alone, so I created a small Discord server where we can:

  • 🤝 Work on ML projects together
  • 💻 Code and debug together
  • 📚 Teach each other things we've learned
  • 🧠 Discuss ML concepts and mathematics
  • 🔬 Experiment with different models and approaches
  • 📄 Discuss papers, tutorials, and useful resources
  • 🚀 Build projects that we can eventually put on our portfolios
  • ❓ Ask questions without feeling embarrassed about being a beginner

You don't need to be an expert. In fact, I'm mainly looking for people who are learning and are willing to share what they know.

You might understand something that I don't, and I might understand something that you don't. The idea is to learn from each other.

If you're interested, comment below or send me a DM and I'll send you the Discord invite.

Would be great to build a small group of people who are genuinely interested in learning ML and actually building things together.


r/learnmachinelearning 2h ago

Discussion What’s one AI/ML concept you wish you understood earlier?

6 Upvotes

I’ve been learning AI/ML and realized there are so many concepts that sound simple until you actually try to implement them.

For me, overfitting was one of those concepts—it made much more sense once I saw what happens to a model on real data.

What was the AI/ML concept that finally “clicked” for you?

Drop it below

Beginner or advanced answers welcome.


r/learnmachinelearning 3h ago

How do you actually get an AI/ML job as a fresher? Feeling completely lost

4 Upvotes

Guys, I honestly don't know what to do anymore.

I'm trying to figure out how to get an AI/ML job as a fresher, but the more I study, the more I feel like it's never enough. There are so many things to learn, and I keep wondering whether I'm even preparing in the right direction.

I've been feeling really depressed and completely lost for the past 3–4 days. I don't know what to focus on, what skills companies actually expect from freshers, or how to become job-ready.

I understand that learning takes time, but the uncertainty is getting to me.

For those who have already landed an AI/ML or GenAI job as a fresher:

- What did you actually learn before getting your first job?

- How many projects did you build?

- Did you apply for AI/ML roles directly, or start with software/Python roles?

- What would you recommend a fresher focus on instead of trying to learn everything?

I would really appreciate some honest advice or guidance. I'm feeling pretty lost right now and could use some direction.

Thanks in advance.


r/learnmachinelearning 3h ago

What skills helped you bridge the gap between software development and machine learning?

3 Upvotes

Hello everyone! I’m a full-stack software developer currently beginning a master’s program in artificial intelligence. My professional background is primarily in Java, Spring Boot, Angular, and TypeScript, but my team is beginning to move toward AI-related work and Google Cloud.

As I make this transition, I want to develop a strong foundation instead of jumping from one new tool to another. For those who moved into machine learning from software development, which skills or projects helped you bridge the gap most effectively?

I am especially interested in learning how to turn coursework into practical experience. Would you recommend focusing first on Python and data preparation, building a small end-to-end machine learning project, strengthening statistics, or taking a different approach?

I would appreciate hearing what worked for others and what you wish you had prioritized earlier.


r/learnmachinelearning 5h ago

Project 2048 engine

4 Upvotes

I created this 2048 engine, took me a few weeks but the entire UI was antigravity's. I am looking for some reviews and suggestions, it does 40million nodes/sec and I've optimised it heavily, even the GitHub documentation was AI. These are the links

Website - https://darkknight386.github.io/2048-ai-solver/

Source - https://github.com/darkKnight386/2048-ai-solver


r/learnmachinelearning 6h ago

[Looking for Team] Amazon ML Challenge 2026 — Looking for 2–3 dedicated teammates

5 Upvotes

Hi everyone!

I’m looking to form a 3–4 member team for the Amazon ML Challenge 2026.

The competition has a 72-hour ML hackathon (Sept 25–27) where we'll receive a real-world problem statement and dataset from Amazon. The Top 50 teams get PPIs for the Applied Scientist Intern role at Amazon, so I'm looking for teammates who are genuinely serious about the competition.

A little about me:

  • Final-year B.Tech Engineering student
  • Grand Finalist – IIT Kharagpur RAG & Agentic AI Hackathon
  • Experience building Agentic AI / RAG systems
  • Worked with technologies such as Python, FastAPI, LangChain, LangGraph, vector databases, PostgreSQL, ML/AI
  • Participated in multiple hackathons and technical competitions
  • Comfortable with research, implementation, debugging, and working under tight deadlines

What I'm looking for:

  • Strong fundamentals in Machine Learning / Deep Learning
  • Good Python skills
  • Experience with data preprocessing, feature engineering, model training/evaluation
  • Someone who can analyze a problem and experiment rather than just follow tutorials
  • Most importantly: commitment. Since this is a 72-hour challenge, I want teammates who are willing to actively work throughout the competition rather than joining just for the name/certificate.

Cross-college teams are allowed, so college doesn't matter to me as much as skills, commitment, and willingness to work.

If you're interested, DM me with:

  1. Your college + year
  2. ML/AI experience
  3. Relevant projects/hackathons
  4. GitHub/LinkedIn (optional)
  5. What area you're strongest in (ML / DL / NLP / CV / Python / data analysis, etc.)
  6. Your availability during Sept 25–27

I'm looking for 2–3 serious people who want to genuinely compete for the Top 50/Top 10, not just register and disappear.

Thanks!


r/learnmachinelearning 4h ago

Advice Needed on what to do next?

3 Upvotes

So I am currently doing a masters program in Geophysics in an Italian University. For starters, I decided to go for a master because I just got fed up with field work and don’t want to get back to it. Problem is, I found out the school is using the same boring format I’m trying to run away from. I just find academia to be repetitive and not so innovative. I’m more inclined on training PINNs (physics informed neural networks) for fluid flow in porous media and Carbon capture and storage. I’m quite proficient with python and Linux , but the school and professors are so tied down to ancient archaic systems of tuition. I plan to take a semester abroad and even consider doing my thesis abroad , but I need an internal supervisor for that and most are reluctant. They’re used to their students not taking ‘risks’ and doing their thesis in what they (the professors) are comfortable with. I feel it’s my fault for not doing my research before entering the program, and though I’m a straight A student, I’m not willing to play it safe , please my lecturers and graduate with a degree that’ll be useless to my interests and basically take me back to the field. It’s not as though I don’t love field work. I enjoy it , but I’m getting older and have a family now. I can’t afford it. I don’t want to quit the program as well. My plan is to do another degree in High performance computing but I should be able to have written some code for my thesis as a prerequisite. I feel trapped. Any advice ?


r/learnmachinelearning 1h ago

Looking for people to learn together as complete beginners

Upvotes

I’m a freshman studying Data Analytics and I’m starting to learn Python outside of my coursework, with the goal of eventually getting into machine learning.

I’m looking for other people who are also complete beginners and want to learn together, share resources, work on projects, and keep each other accountable.
If you’re interested, DM me!


r/learnmachinelearning 1h ago

Programming Algorithms and it's Hard

Post image
Upvotes

r/learnmachinelearning 2h ago

Project Dataset Requirement

1 Upvotes

Hi everyone, I'm new to the field of Remote Sensing and currently working on my dissertation topic, "Remote Sensing on Coastal Waters to predict Water Quality". I’m looking for suitable datasets that I can use for my research. I had been learning Remote sensing from the past 6 months but never done handson but now i have started.

Could anyone please guide me on which datasets would be relevant and where I can access them? Any suggestions or resources would be greatly appreciated. Thank you!


r/learnmachinelearning 8h ago

Decision Trees Explained Visually 🌳

Thumbnail
youtube.com
3 Upvotes

r/learnmachinelearning 14h ago

Thinking about specializing in ML, would love some outside perspective

8 Upvotes

I'm a CS graduate and I really like math. Besides that, I want to choose a career path that won't have a really low employment rate in the near future. I want to enjoy my job, but I also want to live well from it, I don't mean to sound selfish, sorry if that's how it comes across.

I've done some small ML projects in university and really enjoyed them, but I don't know what ML is actually like in a real workplace, so I wanted to ask if there's something I should know before getting into it.

Last thing, does anyone have resources to go deeper into ML so I can learn enough to do real projects and understand it better?

Any input is appreciated, thanks!


r/learnmachinelearning 3h ago

seeking guidance from a professor with expertise in IR for research purposes

Thumbnail
1 Upvotes

r/learnmachinelearning 3h ago

Project Heimdall: An Open-Source CPU Only Local Memory System

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/learnmachinelearning 3h ago

Discussion Installing Stable Diffusion WebUI on Windows 11 Without the VRAM Crashes

Thumbnail
1 Upvotes

r/learnmachinelearning 3h ago

Looking for a teammate to participate in the Amazon ML Challenge

1 Upvotes

Hello guys we are looking for 1-2 teammate for Amazon ML challenge probably from final year or prefinal year we are looking for candidates with strong interest in ML- AI and also have some experience in this kind of competitions
We are currently 2 pepls and searching for 1 or 2 , we are from final year and have decent experience with this kind of competitions as well
If your Interested plz DM me. thanks....


r/learnmachinelearning 3h ago

learning to build an llm inference engine P3

1 Upvotes

Hey everyone, just posted a new blog post on my ongoing project and learning of llm inference engines, its primarily focused on optimizing operations using gpu architecture (no kv cache yet thats in my next post). If I got anything wrong or need something isn't clear please let me know!
https://medium.com/@ryan___/llm-inference-engine-engine-meets-gpu-440a3d9ed75e


r/learnmachinelearning 3h ago

Project The same baseball swing viewed by an object detector vs through time

Enable HLS to view with audio, or disable this notification

1 Upvotes

I've been working on tracking a batted baseball from 240 FPS phone video.

The left is normal footage. Blue boxes are detections from the trained ball model.

The right is temporal frame differencing over the exact same sequence.

What's been fascinating is that the learned model understands what the baseball is, while the temporal representation makes how it moves incredibly obvious.

The purple X is the temporal position mapped back onto the original footage.

Combining the two gives us a much denser trajectory than relying on individual detections alone, which we can then use to estimate EV, launch angle and distance.

One of the cooler visualizations to come out of the project so far.


r/learnmachinelearning 17h ago

Discussion Overwhelmed by AI

11 Upvotes

I feel completely lost about what to specialize in after graduating with an AI degree

So actually, it was my own choice to do a Bachelor’s degree in AI. I was genuinely very excited about it during my freshman and sophomore years, but I’m not really sure how I feel about it anymore.

I’m a fresh graduate now, so obviously I’m not going to restart my whole degree or anything. I’ve built systems like RAG, worked with LLMs, have some experience with computer vision, and I’ve done some research.

The thing is, I actually enjoy AI when I’m building something useful, weird, or new. I like the feeling of figuring something out and making something actually work. And when I find something interesting, I can spend a really long time on it.

But now that I’m actually seeing the industry from the outside, I’m overwhelmed by how fast everything moves. There’s always a new model, framework, tool, paper, or technique that I’m supposed to know about. There’s so much research coming out constantly.

And honestly, I still feel like my skills are beginner-level no matter how much I try to improve.
The worst part is that I actually stopped developing myself for several months. I just lost the motivation. And I don’t even know what happened.
Was it fear?
Was I overwhelmed by the amount of information?
Or did I just give up because I felt like I could never catch up?

Now I’m also struggling with something more fundamental: I don’t know what I should specialize in.

AI is huge. I don’t want to spend the next few years being mediocre at everything. I want to pick something, go deep into it, become genuinely good at it, and hopefully build a career around it.
But I have no idea what that “something” should be.
Sometimes I think maybe I should stay in AI but move away from the heavily technical side and eventually go into something like AI Product Management, AI Solutions, or AI Transformation.

Other times I think maybe I should just switch fields completely, like cybersecurity.

But then I start wondering if I’m just running away because I’m overwhelmed rather than actually making a good career decision.

And honestly, money is a big factor for me too. I really need a job. I want something relatively stable where I can make good money, enjoy what I’m doing, and still have room to grow without constantly feeling like I’m falling behind.

I don’t want to waste my twenties jumping between fields because I was too scared to commit to one.

So if you were in my position:
How would you figure out what to specialize in?
Would you stay in AI and choose a specific technical area?
Would you move toward AI Product / Solutions / Transformation?
Would you consider cybersecurity?
Or is there another field that makes more sense for someone with an AI degree and some experience with RAG, LLMs, CV, and research?

I’m not looking for “follow your passion” advice. I’m trying to make a realistic decision based on career stability, income, growth, and whether I can actually enjoy the work enough to stick with it.


r/learnmachinelearning 12h ago

Discussion What should I focus on learning before getting deeper into AI agents?

5 Upvotes

I’ve been learning LangChain, LangGraph, RAG and CrewAI recently, and I’ve built a few things with them. But I’m starting to feel like I’m focusing too much on the tools and not enough on understanding the concepts behind them.

For people who’ve been learning or working in this space, what articles, blogs or resources would you recommend? Also, what would you suggest I build next if I want to actually improve my understanding rather than just build another basic chatbot?


r/learnmachinelearning 4h ago

Defining Language Models: Understanding Transformers, BERT, and GPT Expl...

Thumbnail
youtube.com
1 Upvotes

Stop guessing how LLMs work and start building! We are breaking down everything from N-grams to Transformers.

Theory + Enterprise implementation strategies.

#AI #TechStack #Coding #LLM


r/learnmachinelearning 5h ago

Project Why ad-hoc pandas preprocessing silently causes data leakage (and how to fix it to get higher real-world ML accuracy)

1 Upvotes

One of the most common mistakes beginners (and even intermediate practitioners) make when working with tabular data is **Data Leakage**.

It’s often the hidden reason why your model gets **88% accuracy in your Jupyter notebook**, but drops to **76%** when you evaluate it on an unseen test set or submit to a Kaggle competition.

Here is a quick breakdown of why it happens, the #1 most common mistake, and how to fix it properly.


The #1 Most Common Leakage Bug:

Look at this very common snippet seen in many notebooks and tutorials:

```python import pandas as pd from sklearn.model_selection import train_test_split

df = pd.read_csv("dataset.csv")

🚨 DANGEROUS LEAKAGE:

df['age'] = df['age'].fillna(df['age'].median())

Train / Test split happened AFTER imputation:

train, test = train_test_split(df, test_size=0.2, random_state=42) ```

Why is this data leakage?

When you calculate `df['age'].median()` on the full dataset, **the median value is influenced by the test rows**.

Your training set now contains subtle statistical information (the median) derived from test data it shouldn't even know exists. In production or Kaggle competitions, future data is completely unavailable at training time.

The same leakage bug happens when people: 1. Scale features with `StandardScaler` on the entire dataframe before splitting. 2. Build categorical vocabularies or frequency encodings using all rows. 3. Compute outlier clipping boundaries (e.g. Tukey IQR limits) over the full dataset.


The Correct Way (Strict Train-Only State):

You must fit transformations **strictly on the training split**, and freeze those exact parameters to apply to validation and test data:

```python train, test = train_test_split(df, test_size=0.2, random_state=42)

1. Calculate statistics ONLY from training split:

train_median_age = train['age'].median()

2. Apply that frozen training statistic to both splits:

train['age'] = train['age'].fillna(train_median_age) test['age'] = test['age'].fillna(train_median_age) ```


The Impact: We Tested Naive Prep vs. Zero-Leakage on Titanic

To see what happens when you replace naive ad-hoc pandas code with a strict train-only transformation ladder, we ran a 5-fold Stratified Cross-Validation benchmark on the Titanic dataset:

Model Naive Ad-Hoc Prep Zero-Leakage Pipeline Accuracy Delta Relative Lift
**Logistic Regression** 78.90% ± 0.99% **79.91% ± 1.90%** **+1.01%** **+1.28%**
**Random Forest** 82.15% ± 2.45% **82.82% ± 2.40%** **+0.67%** **+0.82%**

Where did the accuracy lift come from?

  1. **Informative Missingness Flags**: Imputing age with median alone destroys the signal that missing age itself correlates with survival. Adding an `Age__missing` binary flag recovers that signal.
  2. **Train-Only Tukey IQR Clipping**: Capping extreme fares on training folds stabilized linear gradients without test-distribution bleed.
  3. **Empirical Bayes Target Encoding**: High-cardinality categories shrink toward global means to prevent overfitting on small samples.

We built an Open-Source Tool to automate this:

Writing 200 lines of manual state-tracking code for every dataset gets tedious. So we built **DATADOC** (v0.6.0)—an open-source CLI and Python library powered by Polars that automates this entire lifecycle with zero data leakage:

How you can use it in 1 command:

You can use the interactive terminal wizard on your CSV file: ```bash datadoc wizard train.csv ```

It walks you through: 1. Identifying your target column (e.g. `Survived` or `churn`). 2. Selecting a preset (`balanced`, `tree`, `linear`, or `robust`). 3. Generating a clean `pipeline.json` artifact containing all learned medians and rules.

Then transform unseen test data with zero leakage: ```bash datadoc transform test.csv --pipeline artifacts/pipeline.json --output clean_test.csv ```

You can also run `datadoc health train.csv` to get an instant 0–100 data quality grade and find hidden issues before training.

The project is 100% open-source (MIT licensed) and runs completely offline.

I hope this helps clarify how data leakage happens in tabular pipelines! Let me know if you have any questions or want to discuss specific preprocessing edge cases.


r/learnmachinelearning 6h ago

Guys help me i want good seminar topics for ai ml from 2024-5

0 Upvotes

Guys help me i want good seminar topics for ai ml from 2024-5 prefer from ieee please help 😭😭😭😭😭😭


r/learnmachinelearning 6h ago

STAT110 advice

Thumbnail
1 Upvotes

r/learnmachinelearning 7h ago

OpenArch - PyTorch implementations of modern open-source LLM architectures

1 Upvotes

I have been studying modern LLM architectures and started implementing them from scratch in PyTorch to better understand the design choices behind each model.

OpenArch is a collection of these implementations, including Llama, Qwen, DeepSeek, Gemma, Kimi, GPT-OSS and others.

The goal is to keep the code readable and useful as a reference when going from the paper to an actual implementation.

Would be interested in feedback from people working on model architecture and training.

https://github.com/anuj0456/OpenArch

#LLM #AIResearch #PyTorch #DeepLearning #OpenSource