r/genomics • u/HungarySam • 17h ago
r/genomics • u/three_martini_lunch • Aug 22 '25
New moderator of r/genomics
Hi all
I am taking over the sub as moderator. I am cleaning up stock pumping, spam and other low quality or questionable content.
Please note the new rules aimed at high quality content related to the scientific discipline of genomics.
Please flag posts that do not follow the rules. I am open to additional rules or clarification of the the rules.
r/genomics • u/Dizzy_Upstairs_7581 • 1d ago
We built a queryable knowledge graph connecting 1.1M microbial taxa to diseases, metabolites, pathways, and drugs — sign up for the API
Hey r/genomics,
We've been working on a project called MicroMap — a knowledge graph that integrates microbiome-related data from multiple public databases into a single queryable resource. Wanted to share it here since this is the kind of thing we wished existed when we started doing microbiome research.
What's in it:
- 1,101,289 microbial taxa (NCBI Taxonomy)
- 1,464 human diseases with microbiome associations (Disbiome, BugSigDB, gutMDisorder)
- 6,534 metabolites (HMDB) and 231,556 taxon-metabolite production relationships
- 1,710 metabolic pathways (KEGG, Reactome)
- 6,220 drugs and 1,659 protein targets (ChEMBL)
- 276,169 antimicrobial resistance links (CARD)
- 10,000+ scientific papers with entity cross-references
What you can do with it:
- Query taxa-disease associations with provenance (which paper, which study, what direction)
- Find metabolites produced by a given taxon, or taxa that produce a given metabolite
- Traverse shortest paths between any two entities (e.g., "how is Akkermansia muciniphila connected to Type 2 Diabetes?")
- Identify biomarker signatures and probiotic candidates for a given condition
- Pull cross-feeding networks between microbial communities
Technical details:
Built on Neo4j. The API is RESTful (FastAPI), returns JSON, and supports full-text search across all entity types. Rate limit is 100 requests/minute per API key.
We integrated data from: NCBI Taxonomy, Disbiome, BugSigDB, gutMDisorder, HMDB, KEGG, ChEMBL, Reactome, PubMed, PubChem, and CARD. One of the hardest parts was entity reconciliation — the same organism can appear under different names, different taxonomic ranks, or outdated nomenclature across these sources. Happy to talk about how we handled that if anyone's interested.
Access: https://graphomics.com - email us to get access!
This is part of a broader platform we're building at Graphomics (AI tools for life sciences research), but MicroMap stands on its own as a resource. We'd genuinely love feedback from this community — what data sources are we missing? What queries would be useful that we haven't thought of?
Happy to answer any questions about the data, the architecture, or the integration process.
r/genomics • u/HungarySam • 3d ago
The nuclear pore complex (7R5J & 7R5K) – one of the largest molecular machines in biology.
Enable HLS to view with audio, or disable this notification
The nuclear pore complex is the gatekeeper of the cell nucleus – controlling everything that enters or leaves. We captured it in two states: open (dilated) and closed (constricted).
961665 quantum records
EXXOGENThe nuclear pore complex is the gatekeeper of the cell nucleus – controlling everything that enters or leaves. We captured it in two states: open (dilated) and closed (constricted).
961665 quantum records
r/genomics • u/HungarySam • 4d ago
EXXOGEN - Mapped the full TITIN quantum interactons
galleryTITIN csv 1M+ analytic quantum interaction records
r/genomics • u/BruceNeverWins • 3d ago
Genomic Data Aggregator Core Asset
sideprojectors.comr/genomics • u/berkcat • 4d ago
Visualizing phylogenetic conflict across genomic windows
Author here—I am the first author of this paper. We developed Phylo-Movies because conventional tree-distance measures show how much neighboring trees differ, but not which taxa or subtrees changed position. The paper demonstrates the method using a norovirus recombination boundary and rogue taxa across bootstrap trees. The software and browser demonstration are freely available. https://enesberksakalli.github.io/phylo-movies/ https://academic.oup.com/mbe/article/43/8/msag194/8759530
r/genomics • u/chad_storer • 6d ago
CompBio/MIRaS: Beyond pathway enrichment, a new kind of ‘omics AI
Several years ago, our group saw a need to create a tool that mirrors expert scientists’ ability to look across a messy set of genes, proteins, or metabolites and recognize the biological processes that are contextually enriched based on what they know.
The problem was that human reasoning is powerful, but slow, subjective, and limited to the amount of information a single person can possibly hold.
CompBio/MIRaS takes a different approach from LLMs or pathway enrichment tools by employing methods that unexpectedly converged with theories of hippocampal memory formation, storage, and retrieval. MIRaS is a memory-based associative reasoning engine that explicitly stores biological knowledge as memories, reasons across their relationships, and forms new semantic knowledge through inference. CompBio turns those results into an interactive, traceable map of the biology in your dataset.
Importantly, this analysis is not dependent on matching your dataset with canonical pathways, other datasets, or predefined gene sets. All associations are created from the literature memories identified by your input list, creating low redundancy and contextually relevant results that are fully traceable. Additionally, CompBio includes tools for large scale comparison of knowledge maps, allowing identification of conserved biological patterns across samples, conditions, projects, or reference datasets.
After years of use at WashU and with collaborators, CompBio/MIRaS is now described in our new Nucleic Acids Research paper and is freely available to academic and non-profit researchers.
https://academic.oup.com/nar/article/54/16/gkag833/8769250
If you work with transcriptomics, proteomics, metabolomics, or other complex biological data and this sounds different enough to make you curious, DM me and I can help you get free access.
r/genomics • u/Ok_Fun_3768 • 7d ago
I built a free, comprehensive tutorial site for scRNA-seq, HPC, and Bioinformatics (Scanpy & Seurat)
The Omics Hub is a free learning resource for people starting with computational genomics and scRNA-seq workflows: https://theomicshub.com/
It is designed for learners with biology experience who are new to the command line, HPC environments, and analysis steps such as QC, normalization, clustering, and interpretation. It includes examples in R/Seurat and Python/Scanpy to provide a structured route into genomic-data analysis.
This is my own work, designed from my notebooks, notes, practical workflow experience, and skills. I used AI only to assist with organizing or drafting some sections, while retaining authorship and technical review. I welcome specific technical feedback on missing references, unclear assumptions, version-sensitive steps, or concepts that need clearer explanation.
r/genomics • u/dr3dx • 9d ago
Cpt. T-Cell is a little bit cocky today
When you’ve got a perfectly folded T-cell receptor, a high-affinity match on the MHC-I complex, and a fresh payload of perforin, humility tends to take a backseat.
He’s probably strutting through the lymphatic highways, flexing his CD8 co-receptor, and demanding every cell show its molecular ID. One suspicious non-self peptide, and he's handing out apoptosis notices without a second thought. You can hardly blame him; floating around with that level of precise cytotoxic authority goes straight to a cell’s nucleus.
Did he just successfully eliminate a major viral threat, or is he throwing his weight around over a harmless bit of pollen?
r/genomics • u/THBAX20 • 10d ago
Protein Structure and Sequence Annotation Tool
Hello, I've developed a web-based platform called AlphaSuite Atlas that automatically annotates protein structures with their functional regions in seconds.
You can search over 570,000 proteins and over 11 million structures by name, species, UniProt ID, PDB code, disease, pathway, or plain English (e.g. DNA binding proteins involved in breast cancer).
In around 15 seconds, Atlas returns fully annotated, interactive structure and sequence, mapped with functional domains, motifs, secondary structure, ligands, cofactors, and a plain-language summary of what each component actually does.
Every available structure for a protein (both experimental and predicted) can be accessed and uniformly annotated, with links back to the original papers and databases so all the underlying resources are right there.
Its not finished and we have some bugs to work out so I'd love to hear any feedback after you give it a try here: https://alphasuite.bio/waitlist
Heres a survey to give feedback: https://forms.gle/BEHLNoHgjbLqnSEj7 But feel free to message/email with any further feedback or questions.
Looking forward to hearing your thoughts :)P
r/genomics • u/Specialist-Tune-4158 • 11d ago
Need help in using cellranger with sgRNA/CRISPR sample/Purtub seq
Before post the question, I figured that some context is needed.
Here is the study: We have human patient samples which we transfected with 1,000 sgRNAs (these sgRNAs are for one gene only, let's call that gene 'X'). Then, the sample was treated with antibiotics to make sure that we select all the cells successfully transfected with sgRNAs. Then, the sample was subjected to scRNA-seq library prep with Chromium Next GEM Single Cell 5' Reagent Kits v2 (Dual Index) with Feature Barcode technology for CRISPR Screening. From the exact same sample, a GEX library was made and a single-cell sgRNA library was made. So in the end, I got two sets of FASTQs: a) For GEX, which worked with Cell Ranger, but I am struggling with b) which was made from sgRNA.
I know that I have to put in details like this in the config file:
fastqs,sample,library_type
/path/to/fastqs,GEX_Sample_Name,Gene Expression
/path/to/fastqs,sgRNA_Sample_Name,CRISPR Guide Capture
But when I do that for all 1,000 sgRNAs, it throws an error saying Cell Ranger cannot work with an sgRNA sequence which is like this, e.g.: ATCGCTAGCTc (it throws an error). Even if I make it uppercase, it's bound to clash with some other sgRNA.
I know I am bound to get trolled for not asking a chatbot, but I thought a genuine answer from this community is much better. Thanks.
r/genomics • u/Historical-File-1215 • 11d ago
What if Neoantigen Discovery Became a Foundation-Model Problem?
r/genomics • u/Future_Issue_6150 • 13d ago
Seeking a lineage-resolved single-cell dataset for a peer-reviewed study of clonal identity across perturbations
jacekhoffman.substack.comr/genomics • u/Slow-Log-3756 • 13d ago
sequencing machines cost and efficacy
Hi,
I am looking for sequencing machines that can do full genome sequencing for dogs. My budget is 30K. I am also looking for something that can do the sequencing quickly (1-3 days).
I would prefer a small device that I can carry to places, but it is not necessary.
Please let me know.
r/genomics • u/GolfAltruistic3230 • 16d ago
Mitochondrial Eve: A Genetic Thread Through Time @EnteMicrobialWorld # #...
youtube.comr/genomics • u/literanista • 16d ago
I used Promethease report to benchmark against Fibromyalgia genetic research
r/genomics • u/Next-Possession-2984 • 19d ago
Looking for testers: Annostat, an open-source CLI for bacterial genome annotation QC and analysis
r/genomics • u/Holodoxa • 19d ago
Forensic investigative genetic genealogy match rate estimated from a nation-wide population register
biorxiv.orgr/genomics • u/obllak • 22d ago
Sequencing.com / bioinformatics review wait time
Has anyone here had their results escalated to sequencing.com’s bioinformatics team for manual review? If so, how long did it actually take to hear back? We were told 3 to 4 days, and we have now been waiting 8 days with no meaningful update.
This is regarding an unexpected, very serious genetic finding in our 14 m/o daughter. The variant was called from 7 out of 36 reads (29 reference reads and 7 alternate reads), which is one of the reasons we desperately want an experienced bioinformatician to look at the raw sequencing data and tell us how confident they are that this is a real constitutional variant.
When we first saw this result, our entire family was devastated. We cried in despair. We barely slept. We have spent the past week frightened, depressed, and obsessively trying to understand what this could mean for our little girl's future.
When you are waiting to find out whether your baby may have a serious genetic condition, every additional day feels unbelievably long.
r/genomics • u/Holodoxa • 22d ago
The Pan-European Impact of the Balkan Hunter-Gatherers
doi.orgr/genomics • u/No_Lion_3319 • 23d ago
DIY WGS analysis using Python/AI. Is it doable, and what’s the best EU provider under €200?
Hi everyone,
I'm a physicist with a background in data analysis (mostly Python), but little to no knowledge of genomics and bioinformatics. Out of pure curiosity, I'm thinking about getting my DNA sequenced.
My plan is to buy a WGS test, download the raw data, and write Python scripts with AI assistance to query my data. Claude seems very optimistic about how doable this is, but before spending my money, I want a reality check from people who actually work with genomic data.
Here is what I want to achieve:
As a power athlete I want to check:ACTN3, ACE, MCT1 / SLC16A1, COL5A1.
Ethnicity: Get an ethnic breakdown.
Future Proof: Whenever new studies come out in the next years, I can just write a quick script to check my existing .vcf file.
My questions for you:
Is this actually doable for a non-bioinformatician? Is querying a .vcf file using Python + AI as straightforward as it sounds, or am I underestimating bioinformatics pipelines (file sizes, reference genomes, formatting issues)?
Which provider do you suggest in Europe? I’m based in Italy.
Is a budget of ~€200 realistic? I’m willing to wait for Black Friday / flash sales.
Thanks in advance for any insights!