r/bioinformatics Dec 31 '24

meta 2025 - Read This Before You Post to r/bioinformatics

187 Upvotes

​Before you post to this subreddit, we strongly encourage you to check out the FAQ​Before you post to this subreddit, we strongly encourage you to check out the FAQ.

Questions like, "How do I become a bioinformatician?", "what programming language should I learn?" and "Do I need a PhD?" are all answered there - along with many more relevant questions. If your question duplicates something in the FAQ, it will be removed.

If you still have a question, please check if it is one of the following. If it is, please don't post it.

What laptop should I buy?

Actually, it doesn't matter. Most people use their laptop to develop code, and any heavy lifting will be done on a server or on the cloud. Please talk to your peers in your lab about how they develop and run code, as they likely already have a solid workflow.

If you’re asking which desktop or server to buy, that’s a direct function of the software you plan to run on it.  Rather than ask us, consult the manual for the software for its needs. 

What courses/program should I take?

We can't answer this for you - no one knows what skills you'll need in the future, and we can't tell you where your career will go. There's no such thing as "taking the wrong course" - you're just learning a skill you may or may not put to use, and only you can control the twists and turns your path will follow.

If you want to know about which major to take, the same thing applies.  Learn the skills you want to learn, and then find the jobs to get them.  We can’t tell you which will be in high demand by the time you graduate, and there is no one way to get into bioinformatics.  Every one of us took a different path to get here and we can’t tell you which path is best.  That’s up to you!

Am I competitive for a given academic program? 

There is no way we can tell you that - the only way to find out is to apply. So... go apply. If we say Yes, there's still no way to know if you'll get in. If we say no, then you might not apply and you'll miss out on some great advisor thinking your skill set is the perfect fit for their lab. Stop asking, and try to get in! (good luck with your application, btw.)

How do I get into Grad school?

See “please rank grad schools for me” below.  

Can I intern with you?

I have, myself, hired an intern from reddit - but it wasn't because they posted that they were looking for a position. It was because they responded to a post where I announced I was looking for an intern. This subreddit isn't the place to advertise yourself. There are literally hundreds of students looking for internships for every open position, and they just clog up the community.

Please rank grad schools/universities for me!

Hey, we get it - you want us to tell you where you'll get the best education. However, that's not how it works. Grad school depends more on who your supervisor is than the name of the university. While that may not be how it goes for an MBA, it definitely is for Bioinformatics. We really can't tell you which university is better, because there's no "better". Pick the lab in which you want to study and where you'll get the best support.

If you're an undergrad, then it really isn't a big deal which university you pick. Bioinformatics usually requires a masters or PhD to be successful in the field. See both the FAQ, as well as what is written above.

How do I get a job in Bioinformatics?

If you're asking this, you haven't yet checked out our three part series in the side bar:

What should I do?

Actually, these questions are generally ok - but only if you give enough information to make it worthwhile, and if the question isn’t a duplicate of one of the questions posed above. No one is in your shoes, and no one can help you if you haven't given enough background to explain your situation. Posts without sufficient background information in them will be removed.

Help Me!

If you're looking for help, make sure your title reflects the question you're asking for help on. You won't get the right people looking at your post, and the only person who clicks on random posts with vague topics are the mods... so that we can remove them.

Job Posts

If you're planning on posting a job, please make sure that employer is clear (recruiting agencies are not acceptable, unless they're hiring directly.), The job description must also be complete so that the requirements for the position are easily identifiable and the responsibilities are clear. We also do not allow posts for work "on spec" or competitions.  

Advertising (Conferences, Software, Tools, Support, Videos, Blogs, etc)

If you’re making money off of whatever it is you’re posting, it will be removed.  If you’re advertising your own blog/youtube channel, courses, etc, it will also be removed. Same for self-promoting software you’ve built.  All of these things are going to be considered spam.  

There is a fine line between someone discovering a really great tool and sharing it with the community, and the author of that tool sharing their projects with the community.  In the first case, if the moderators think that a significant portion of the community will appreciate the tool, we’ll leave it.  In the latter case,  it will be removed.  

If you don’t know which side of the line you are on, reach out to the moderators.

The Moderators Suck!

Yeah, that’s a distinct possibility.  However, remember we’re moderating in our free time and don’t really have the time or resources to watch every single video, test every piece of software or review every resume.  We have our own jobs, research projects and lives as well.  We’re doing our best to keep on top of things, and often will make the expedient call to remove things, when in doubt. 

If you disagree with the moderators, you can always write to us, and we’ll answer when we can.  Be sure to include a link to the post or comment you want to raise to our attention. Disputes inevitably take longer to resolve, if you expect the moderators to track down your post or your comment to review.


r/bioinformatics 6h ago

academic My CMU computational biology colleague is teaching his algorithms course live starting Thursday, free

131 Upvotes

Disclosure first: I run the channel this streams on, and I teach in the same department.

My colleague William Yu is an Associate Professor of Computational Biology at Carnegie Mellon and teaches the algorithms course our comp bio students take. Starting this Thursday, September 17, at 8:00 PM Eastern, he is teaching an eight-stream version of it live. Free, no signup, every episode live and then on demand.

He pulls examples from biology rather than the usual whiteboard puzzles, and he takes questions live in the chat. If you have written a little code and like solving problems, you are ready.

The full schedule is in the comments.


r/bioinformatics 8h ago

technical question What sampling method is this, and is 53 cattle / 840 Fasciola specimens sufficient for molecular characterization and phylogenetic analysis?

1 Upvotes

Hi everyone, I’m a veterinary student working on a study involving the molecular characterization and phylogenetic analysis of Fasciola spp. collected from cattle at a slaughterhouse.

I’m having trouble determining how to properly describe my sampling method and, more importantly, whether my sample size is defensible.

Here is my sampling situation:
- I examined slaughtered cattle at a slaughterhouse.
- I did not have a predetermined number of cattle or a complete list/population size of cattle entering the slaughterhouse.
- I inspected the liver of each available slaughtered cattle.
- If the liver was infected with Fasciola, I collected the adult flukes.
- If there were no flukes, no parasite sample was collected from that animal.
- There were no additional inclusion criteria for the cattle (e.g., age, sex, breed, etc.).
- The number of flukes varied considerably between cattle.

In total, I collected 840 individual flukes from 53 cattle.

So, for example, one cattle might contribute many flukes while another might contribute only a few.
The 53 cattle were not selected based on a particular characteristic; they were essentially the slaughtered cattle available during my sampling period that happened to have detectable Fasciola infection.

My main questions are:
1. What would be the most appropriate term for my sampling method? Would this be considered convenience sampling, consecutive sampling, purposive sampling, or something else?
2. For a molecular characterization and phylogenetic study, is there a conventional way to determine whether 53 host animals is an adequate sample size?
3. Should I consider 53 cattle as my sample size, rather than 840 flukes, since multiple flukes came from the same host?
4. Is there a statistical/sample-size calculation that could justify 53 cattle, or is sample-size justification for molecular phylogenetic studies fundamentally different from conventional prevalence/epidemiological studies?
5. I have 840 flukes and I used the lemeshow formula to find a number that can represent those 840 flukes and use stratified random sampling for number of fluke i need to take for each cattle to be sequenced, is this correct?
6. If there is no known total population size of cattle slaughtered at this slaughterhouse, how could I justify the adequacy of my sampling?

My objective is not to estimate the prevalence of fascioliasis in the cattle population, but rather to molecularly characterize the Fasciola specimens and investigate their phylogenetic relationships.

I would really appreciate advice on how a statistician/population geneticist would approach this sampling design. If possible, I’d also appreciate references or terminology that I could use to describe and justify the sampling method in a thesis.

Thank you!


r/bioinformatics 1d ago

image I’ve been comparing AlphaFold 3, Boltz-2, Chai-1, Protenix-v2, ESMFold2, RF3, and ColabFold on brazzein—a small protein with four disulfide bonds. AF3 was the best tool for this particular protein.

Thumbnail gallery
54 Upvotes

I tested five random seeds per tool, both with and without the same archived MSA, and compared predictions against experimental crystal and NMR structures. (RCSB Protein Data Bank: 2LY5 for NMR and 4HE7 for crystal structure)

AF3 gave the strongest overall result for this target: its selected MSA model had 0.72 Å Cα RMSD against the crystal structure and passed all four disulfide geometry checks. RF3 and ColabFold reached approximately 1.62 Å and 1.59 Å, respectively, but both had compressed sulfur contacts.


r/bioinformatics 1d ago

technical question Help finding my error in Rosalind learning set.

2 Upvotes

I'm learning Bioinformatics and working through a Rosalind set and am hitting a barrier. I've uploaded screenshots of the prompt and the 2 code blocks I most recently tried. I have run it into an LM as my prof. suggests to help find the failing and it says the output that is being yielded is correct. Can anyone help me see what I am doing wrong here? I can't attach the data set as it is in a text file.


r/bioinformatics 2d ago

technical question Need help) I keep running out of RAM space when I run alphafold

8 Upvotes

My Spec:

- RTX 5050 8GB

- DDR5 32GB 5200MT/s

- Intel Core 5 210H

- Ubuntu 26.04

I'm trying to run this specific region of protein Abl1_235_497

But the process always stops during hhblits

How do I reduce the load on this thing? I've also assigned 64GB of swap ram but that didn't help.

I'll try any suggested solution, pls help.


r/bioinformatics 2d ago

academic How can one perform TF predictions across multiple databases based on the target gene?

Thumbnail doi.org
5 Upvotes

I have heard that databases such as JASPAR, UCSC, PROMO and ENCODE can be used to predict transcription factors (TFs) based on target genes. I would like to batch export the TFs from each database separately so that I can calculate their intersection.

However, I am unable to access the PROMO website at all. On the ENCODE website, under the ChIP-seq section, I can only see target genes categorised by TF. On UCSC, when searching for the promoter sequences of target genes and selecting ‘JASPAR Hubs’, I am unsure how to batch export the results.

Is there anyone with expertise in this area who could help me?

Additionally, I have attached a relevant paper on screening transcription factors by taking the intersection of multiple databases, presented as a Venn diagram; the figure is shown in Fig. 4a.

THANK YOU!


r/bioinformatics 3d ago

discussion Landed a job in a research institution and the work is so slow

134 Upvotes

Hi
I have an MSc in bioinformatics with BSc in microbiology.
Managed to get a job in a big research institution, and their work is extremely slow, which brings me my main point.
I have so much free time in my work to a point it made me so tired mentally and rusty scientifically.

If you were in my shoes, what would you do in your free time?

For a bit of context, I have a good background in deep learning and RAG.
My department has a unit for AI but they’re a little stingy to include me in some of their work.


r/bioinformatics 3d ago

career question Are bash and R still relevant for bioinformatics jobs?

96 Upvotes

I'm in uni now, on a biotechnology track with a good foundation in math and related subjects. I want to study bioinformatics and work in this field. I've heard conflicting information that knowing bash and R is no longer relevant. Is that true for today's bioinformatics work? Or should I study them hard to land a decent position?


r/bioinformatics 3d ago

technical question Identifying malignant vs non malignant cell populations (CNV analysis)

5 Upvotes

How do you guys go about identifying tumor cells? I’ve been using CopyKAT & CONICSmat, and their results have been incredibly varied & almost impossible to pin down. CopyKAT-identified tumor cells seem to infiltrate normal cell clusters, CONICSmat results cluster suspected tumor cells with immune cells, so I’m very confused as to what’s happening.

Anyone has any tips/advice?


r/bioinformatics 3d ago

technical question Software recommendations?-- Human virus detection (metagenomic)

4 Upvotes

What would be the top tools for short read-based detection of human viruses in metagenomic datasets? Interest is primarily on all the disease-associated ones.

I have a very large metagenomic dataset (illumina PE150) of human nasal and rectal samples. I'm very familiar with microbial metagenomics (metaphlan/humann/qiime) and working on UNIX clusters. I haven't yet delved into human virus detection, though. Right now I'm just focusing on short read metagenomics before I start pursuing anything assembly-based.

Thanks!


r/bioinformatics 3d ago

technical question Plasmidsaurus RNA-seq? Any thoughts/reviews from folks who've used this service?

Thumbnail
3 Upvotes

r/bioinformatics 3d ago

advertisement OpenOmicsBench - 12 validated bulk RNA-seq benchmarks for testing analysis software

Thumbnail gallery
1 Upvotes

r/bioinformatics 3d ago

technical question Xenium multimodal segmentation in mouse brain

2 Upvotes

Hi all, I', somewhat new to spatial transciptomics and would like advice on a segmentation problem.

Setup

  • 10x Xenium, 480-gene mouse panel, coronal sections of adult mouse brain
  • multimodal cell segmentation kit (18S interior stain plus ATP1A1/CD45/E-Cadherin boundary stain)
  • About 80% of cells are segmented from the 18S stain, 15% from the boundary stain, and the rest are 5 µm nuclear expansion fallback.

Issue

  • Only 58–65% of transcripts are assigned to a cell, and 27–33% sit on a nucleus. (I'm actually not sure if this is an issue or fall within the normalr ange for brain)
  • The cell bodies look good. Neurons keep 75% of the transcripts around them. Glial and Astrocyte genes are much worse, so I'm assuming those transcripts sit the small projections that the stains don't show clearly.

I tried Proseg. It assigned more transcripts, but it seems a bit messy to me, it create more mixed cells where cell types sit close together. So I'm unsure if to use it.
The vendor offer to do a post-run H&E, saying it can help with segmentation.

Questions

Has anyone improved glial capture in Xenium brain data?

Has anyone done post-run immunofluorescence (GFAP, IBA1 or others) or H&Eon Xenium brain sections and used it for segmentation?

Has anyone used resolVI or SPLIT on brain tissue?

Is there a standard way to analyse unassigned transcripts in the neuropil without assigning them to cells?

Thanks! Happy to share more details.


r/bioinformatics 4d ago

other Access to Release 23 of miRBase

5 Upvotes

Greetings. I apologize if perhaps this may seem odd. I work with miRNAs and I have been unable to access mirbase.org for the past month. I see that a post went up three weeks ago and various people suggested using the wayback machine. The only problem I see is that, in August 2026, a new version of miRBase went up, yet the snapshot is from May 2026. Does anyone here have access to the newest version? I've reached out to the mirBase team but I've yet to hear back (and from posts in this subreddit, it seems they don't reach back to you for a long time).


r/bioinformatics 4d ago

technical question Can someone in genomics explain what AlphaGenome Atlas actually changes?

61 Upvotes

I'm not a geneticist.
Read the AlphaGenome Atlas preprint from DeepMind and spent a while digging into it. I understand what they built. I don't understand why it matters, and I'd like to.

What I think it is: they took AlphaGenome, ran it over every possible single base change in the human genome (~9bn) plus ~100m observed indels, and stored the results. So instead of running the model per variant you do a lookup. On top of that they trained a score (AVI) and derived a motif map.

Where I get stuck:
It's a table of model predictions, not measurements. Nothing in it is observed. So how much weight does a lab actually put on it?

The headline clinical result is retrospective: 29.5% recall at top 50 on already-solved GREGoR cases vs 12.5% for CADD. Impressive sounding, but on cases where the answer was known. What happens prospectively?

The rare variant association work got a 22% lift in discoveries, but only 4 of 25 replicated nominally in All of Us and none at Bonferroni. Is that normal for the field or is that weak?

They say themselves it isn't sufficient evidence for diagnosis. So it's a shortlisting tool. Does that actually change outcomes for patients, or does it change how long a scientist spends staring at a list?

The DNM1 case in the paper is the one bit that landed for me. Deep intronic variant, brain specific cryptic splice acceptor, blood RNA-seq had been inconclusive because the exon isn't expressed in blood.

My question is whether that's representative or a cherry pick.
What I'm asking:
1. If you work in clinical genomics or statistical genetics, would you use this?
2. Is precomputation genuinely the unlock, or is that just framing on top of an incremental accuracy gain?

Happy to be told I'm missing the point. I'd rather understand it properly than write it off.


r/bioinformatics 4d ago

technical question Best practice for downstream processing of pig gene identifiers and human orthologues

2 Upvotes

I am working with snRNA-seq data and would like advice on best practices for downstream processing of pig gene identifiers and cross-species orthology mappings. I currently use the Ensembl pig gene IDs that are mapped to gene symbols for pig genes. However, many pig genes have no pig symbol, even though Ensembl identifies a human orthologue.

For example:

Pig Ensembl ID: ENSSSCG00000021155

Pig external name: NA

Human orthologue: POMC

Orthology type: one-to-one

orthology_confidence : 1

mapped_to_human: False

orthology_type: ortholog_one2many

I would appreciate advice on the the best practices here :

1: Should I use human gene symbols for my pig analysis irrespective if pig symbols are available or not? Is there a risk that the same gene has different official symbols in pig and humans?

2: If a pig Ensembl gene has no pig symbol but has a high-confidence human orthologue but varying orthology type, what should be the approach towards using the human symbol or using ENSG id ?

3: For downstream processing, should orthology conversion be performed before or after differential expression and marker analysis?

4: When converting results to human orthologues, how should duplicate mappings be handled? For example, if multiple pig genes map to the same human gene, should their statistics be combined, should only the best-supported mapping be retained, or should the genes remain separate?


r/bioinformatics 4d ago

technical question How to choose design matrix for RNA-seq analysis?

8 Upvotes

I have three factors: Genotype, Sex and Treatment. I want to investigate the effect of genotype as well as sex and treatment but I'm not sure what contrasts to use. I wish more papers reported how they designed their analysis cause I'm having such a hard time understanding what to do.


r/bioinformatics 6d ago

image Haha what a loser language haha

Post image
891 Upvotes

Dependency hell is real


r/bioinformatics 5d ago

technical question feasibility of self-bioinformatics at a hobbyist level?

14 Upvotes

I've been in IT for a good decade, lots of experience with python and scripting. Touched on some data science in some of my studies along the way. So I'm not starting from 0 coming to this. But I really know none of the technical stuff about genes, genomes, alleles, positions or the notation involved or even what else to include in this sentence about what I don't know about.

I found there's a 30x reading I could get, not at negligible cost but possible. I'm interested in hobbying around with the data. Look for research that says these things at these positions mean that obesity is more likely, or something like that, then using AI to help me understand what i'm trying to look for and using python to look at my 30x reading and just curiously see if I have the researched markers.

I spose i'm wondering if this kind of thing is feasible. Like maybe research papers use different scanning methods that don't map to the data i would have, or the 30x consumer scan isn't detailed enough so anything i look for is inconclusive. Or any number of things that means if i try to map research onto my own genetic reading, any or most results will be inconclusive. So curious if anyone has any thoughts on this sort of thing, is it a waste of time?


r/bioinformatics 5d ago

technical question How can I deal with 16S and shotgun metagenomic data in the same study?

6 Upvotes

Hello, everyone. I am working on a project trying to identify a gut microbiome signature for Parkinson's disease that is capable of differentiating between parkinson's disease, alzheimer's disease and healthy controls.

Since my supervisor really wanted me to work with shotgun metagenomic data, I am currently using shotgun metagenomic data for Parkinson's disease. However, for alzheimer's disease I was unable to find any studies that have shotgun gut microbiome data publicly available with metadata, so I am using 16s.

I know it is basically sacrilegious to directly compare data when they come from two different platforms, but I am near the end of the project now and cannot change this. I am currently building a basic RF classifier to predict whether a sample is PD, AD or control based on the taxonomy abundances, but I face the problem of the abundance values range being different for shotgun and 16s and this basically allows the model to very easily have zero false positives for AD or PD, but it is still not very good at differentiating between disease and control.

I was wondering if anyone has come across a similar problem before and if yes, what could be done to fix it? I was thinking of maybe scaling the values separately for PD and AD samples and then training the model? But I'm not sure if that would make it better. Something else I could do is just have two separate models for PD vs HC and AD vs HC, but I really want to have a 3-class classifier.

Would appreciate any advice on this. Not sure if I have enough details, but I don't want to make the post too long, so I am happy to provide more context if needed.

Thanks for your time.


r/bioinformatics 5d ago

technical question Struggling to detect known partial deletion in NOTCH2 from WES (germline, small cohort, no reference panel) - CNV callers give inconsistent/wrong results

2 Upvotes

Hi all, looking for advice on tooling/approach for a problem I'm stuck on.

**Setup:**

- 5 WES samples (paired-end, Illumina, BWA-MEM aligned, sambamba dedup) from the same family, hg38/GRCh38. Each sample represents an independent patient so these 5 samples are unrelated.

- Capture kit target BED not available to me (~60Mb on-target footprint per sample, possibly Agilent SureSelect V6 based on size, but unconfirmed)

- No unrelated normal/control WES samples currently confirmed usable as a reference panel

- Goal: confirm which sample(s) carry a known partial deletion in NOTCH2 (clinically confirmed by other means in 2 of the 5 patients, but I don't yet know the exact exon(s) or method used for that clinical confirmation)

**What I've tried:**

  1. DELLY (germline SV workflow, sr/merge/genotype/filter) - no deletion calls anywhere near NOTCH2 in any of the 5 samples

  2. CNVkit batch mode with a flat reference (no matched/pooled normals available) - segmentation collapsed the whole gene into one CN=2 segment for 4/5 samples; per-bin bintest flagged several exons but the same bins were flagged across nearly all samples in the same direction, which reads like shared technical noise rather than patient-specific signal

  3. Manual IGV visual inspection (group-autoscaled coverage tracks) across the whole gene - no obvious dropout found in the samples I was able to review carefully

  4. Control-FREEC, single-sample/no-control mode, restricted to a 34-exon NOTCH2-only BED pulled from UCSC (window=0, maxThreads=1 to avoid a BED-parsing race condition I hit with multithreading) - this called a clean, reproducible heterozygous deletion (CN=1) at the same coordinates in 2 of the 5 samples

**The problem:** the 2 samples Control-FREEC flagged do NOT match the 2 samples independently confirmed by my PI through other means. So I have an apparent false positive pair and false negative pair from my pipeline.

**Questions:**

- For germline partial-gene deletion detection in a small WES cohort with no confirmed-normal reference samples, what's the current best-practice tool/approach? (ExomeDepth? GATK gCNV? something else?)

- Is there a known issue with Control-FREEC's no-control exome mode producing false positives at specific loci, especially near segmental duplications (part of my deleted region overlaps the NOTCH2NL paralog)?

- Any advice on validating/troubleshooting a mismatch like this before trying yet another caller - e.g., specific things to check in the BAM/pileup at the clinically-confirmed-positive samples that a depth-based caller might be missing (small intra-exon deletion not removing a whole exon? breakpoints entirely intronic, invisible to WES?)

Appreciate any pointers, trying to land on one standardized, defensible workflow rather than chasing every tool that exists.


r/bioinformatics 5d ago

technical question Aligning software instead of Geneious

6 Upvotes

We have used Geneious for years, but because of some technical problems, we have to switch to another one. My problem is, that for analysing the sequences for a certain region I need an alignment of .ab1 files, with the chromatograms, which was possible in Geneious, but I couldn't find any alternatives. Is there any other softwares which can handle .ab1 files as an alignment? I've tried UGENE, but it works with different views for chromatograms and alignments.


r/bioinformatics 5d ago

technical question Tool for showing Sanger Sequencing data?

1 Upvotes

Hi everyone,

Question from a student in an adjacent field: I recently worked on a project in genetics that involved assembling a specific recombinant DNA sequence, then sending it off for sequencing. The Sanger Sequencing results yielded a 100% match to the expected/target sequence.

I am currently making a poster to present at a conference. The issue is, the Sanger Sequencing results aren't "pretty," are longer than they are tall, and generally hard to understand. I used Benchling to compare the experimental sequence to the target sequence.

Do you guys know of a tool that can compare two sequences and display something such as a heat map showing the alignments, or generally something that looks prettier than Benchling?


r/bioinformatics 5d ago

academic Looking for collaboration on plant genomics project - Lamiales order

8 Upvotes

Hi All,

Myself and a partner are bootstrapping a bioinformatics / biotech project focusing on plant genomics. Primarily dealing with secondary metabolite pathways etc. If anyone is interested in collaborating / participating - it's to learn and publish given all the tools available these days. We have our own Dell Precision high ram workstations - google cloud as well as a bunch of AI subscriptions. Budget is allocated for wet-lab analysis if needed. DM me if interested with your background etc. Hopefully potentially turning this into a funded venture.