User:Madzgopher
Problem Set 1
0) AI Overview
The FAM200B is a protein gene (Family With Sequence Similarity 200 Member B),when in humans is encoded by c5orf53, is a protein-coding gene located on chromosome 4p15.32. The human variant transcript (NM_001145191.2) encodes a 657–amino acid protein(UniProt reviewed entry P0CF97). Its physiological function remains uncharacterized, expression data confirm that its RNA is present across many human tissues, including brain and thyroid, with moderate levels (RPKM ~5–6). Protein-level immunohistochemistry has been attempted, but the Human Protein Atlas reports insufficient reliability for tissue-level estimation.
Proteomics data confirm that FAM200B protein exists in vivo, meeting criteria for Protein Evidence Level 1 (PE1) in neXtProt/ProteomeXchange. In functional studies, a large-scale RNA interference screen showed that siRNA-mediated knockdown of FAM200B inhibited HIV-1 replication in TZM-bl cells (Brass et al., Science, 2008), suggesting a potential but uncharacterized role in viral host interactions. 6
No FAM200B associations are established in ClinGen or OMIM, and FAM200B has no confirmed clinical phenotype links. It has a known paralog (FAM200A) and multiple orthologs across vertebrates, indicating evolutionary conservation, but its molecular role remains undefined.
The AI wasn't very helpful. The first paragraph was information I had already known and determined the protein was Level 1 meaning it's real and in humans which I already knew. Lastly it gave me a potential lead about function but didn't give me any solid sources (not sure what Brass 2008 was) I could confirm this information. Overall not very helpful but I am going to keep looking into HIV studies.
- Bioinformation: Accession number: >NM_001145191.21, Protein: >NP_001138663.12
Aliases: C4orf54
Homo sapiens FAM200B is a 4287 nucleotides gene located on chromosome 4 at 4p15.321 and encodes a 657 amino acid protein2. The gene has 2 exons, the second one consisting of the base pairs 76 to 4287 of the mRNAcontaining almost the whole gene1. The gene has 23 transcript variants encoding 2 protein isoforms. FAM200B transcript mRNA is found in multiple cell lines, human tissues and cancer samples however with low tissue specificity in all categories; meaning it's not specific to a particular tissue or region.5 The gene has orthologs in mammals, birds, reptiles, amphibians and fish. FAM 200B has over — paralog with the most important being FAM200A with 79.79 % identity. Omin and wikipedia concluded no search results however PubMed found some scientific articles containing “FAM200B”. In a Mendelian randomization study using dopaminergic neuron–specific eQTL data, altered expression of FAM200B was identified as significantly associated with Parkinson’s disease risk, meaning it's a potential gene for disease susceptibility.3 FAM200 B showed low tissue specificity for the brain however is present in the basal ganglia, the part of the brain most associated with Parkinson's disease. Within the brain the gene is 3rd most in the basal ganglia behind the hypothalamus and white matter.5 In another study looking at childhood constipation, decreased expression of FAM200B was observed in patients and identified as a hub gene associated with reduced CD56⁺ NK cell counts (p < 0.05).4 Both articles suggest the gene is present in many parts of the body and responsible for many possible functions.
Repeat 1: Expressed in my needle pairwise alignment for Homo sapiens and Homo sapiens. Looked for internal structure of FAM200B and found 2 repeats; figure 1 and 2.
- Sequence Bioinformation
> FAM200B - 4289 nucleotides
[nnagaggcgcagctctgctggagagacctgaggcttgcagcgctgggcccggccgggtcg [exon 1]/[exon 2]
tgcggctgggctgga][ggtcaaggatgcattccaggaaagttctataaaaatataccatc
aactccaacctgatatttgacctagccatggaaaacctcattggctcctcactggatcag
accataaaaatgatccaggaatgaaaatctaagataaattcccaccacttcaagatgttc
ttatacaagtcaacagtctagcagatcagccagtctctttcatcatcaagcagacatact
caaggtaccaacacagcaaaggcacaaaatgtctctatatcagtatttttcatggtatat
gtgagacccattagtgaggcatgaaatcaattctgtgagtcacaatgaacatttgtctaa
aaaacaaaataaaatgtaaaataccaaagtgttttgcatgaatcctttgcctgaattgtg
cacattgcaaggaatgctgaggtcaatttcaatatgcgagaggggtatttgatggaggag
atatgatgaatagaatgtatattggctaagaaatggttttcagtcattgattcctgggtg
tcttgcttgattttcataatagcacggttaatgttctatagtcaagttttatatgattacaat
tacttgcaactagttgtgaagtctgaaatattcgattcaattgcaaactttgagtgtagt
ttttgaaaagatgttatttagactgtatatttttttcttcttttttgagttagtgccaat
Tataacattttaatcaaactggaacaaattgctattaaaatggatcatttctttattaaa start codon
Y N I L I K L E Q I A I K M D H F F I K
agaaagaggaatagtgaagtgaaatatacagaagcatgttcaagttcatctgttgaatct
R K R N S E V K Y T E A C S S S S V E S
ggaattgtgaatagtgacaatattgagaaaaatactgactccaatctgcaaacttcaact
G I V N S D N I E K N T D S N L Q T S T
tcatttgagccacatttcaaaaagaaaaaagtaagtgcaagacgttataatgaagattac
S F E P H F K K K K V S A R R Y N E D Y
ttaaaatatggctttatcaaatgtgaaaaaccctttgaaaatgacagacctcagtgtgtt
L K Y G F I K C E K P F E N D R P Q C V
atttgtaataatattcttgcgaatgaaagcttaaaaccttcgaaattaaaaaggcactta
I C N N I L A N E S L K P S K L K R H L repeat 2B
gaaactcagcatgctgaacttattgataagcctcttgaatattttcaaagaaagaaaaaa
E T Q H A E L I D K P L E Y F Q R K K K
gacataaagttatcaacacaatttcttagttgttctactgctgttagtgagaaagcctta
D I K L S T Q F L S C S T A V S E K A L
ttatcatcatatttagttgcatatcgtgtggcaaaagagaaaatagctaacacagctgct
L S S Y L V A Y R V A K E K I A N T A A
gaaaaaattattcttccagcatgtttggatatggtgcgtacaatatttgatgataaatca
E K I I L P A C L D M V R T I F D D K S
gctgataaattaaaaactatacctaatgataacacagtatctcttcgaatttgtactatt
A D K L K T I P N D N T V S L R I C T I repeat 2A
gcagaacatttagaaacaatgcttattactcgtttacagtctggtatagattttgcaatc
A E H L E T M L I T R L Q S G I D F A I
cagcttgatgaaagcactgatattggaagctgcacaacacttttagtttatgtcagatat
Q L D E S T D I G S C T T L L V Y V R Y
gcgtggcaagatgattttttggaggattttttgtgttttttaaatttaacctcacaccta
A W Q D D F L E D F L C F L N L T S H L
agtggattagatatttttacagaattagaaaggcgcatagttggccaatataaattaaac
S G L D I F T E L E R R I V G Q Y K L N
tggaaaaactgtaaaggaattacaagtgatggcacagcaaccatgactggaaaacatagc
W K N C K G I T S D G T A T M T G K H S
agagtaattaaaaaattactagaagttactaataatggtgctgtgtggaatcattgtttt
R V I K K [L L E V T N N G A V W N H C F] repeat 1A
atacatcgtgaaggtttagcatccagagaaattccacagaatctcatggaggtattgaaa
I H R E G L A S R E I P Q N L M E V L K
aatgcagtgaaagttgttaattttattaaaggaagctcattgaatagccggcttcttgaa
N A V K V V N F I K G S S L N S R L L E
acattttgttcagagattggaactaatcatacccacttactatatcataccaaaattcgt
T F C S E I G T N H T H L L Y H T K I R
tggttgtctcaagggaaaatactaagcagggtttatgagctcaggaatgagattcacttt
W L S Q G K I L S R V Y E L R N E I H F
tttctcattgaaaaaaaatctcatttggcaagtatttttgaagatgatacttgggtaaca
F L I E K K S H L A S I F E D D T W V T
aaattggcatatttaactgatatttttagcattcttaatgaactgagtttaaaactacag
K L A Y L T D I F S I L N E L S L K L Q
gggaaaaacagtgatgtattccaacatgttgaacgtatccagggatttcgaaagacatta
G K N S D V F Q H V E R I Q G F R K T L
ttgttatggcaagtaagacttaaaagtaatcgtcctagctattacatgtttccaagattt
L L W Q V R L K S N R P S Y Y M F P R F
ttgcagcatattgaagagaatattattaatgaaaacattttgaaagaaataaaattagag
L Q H I E E N I I N E N I L K E I K L E
atattgttgcatctcacttctctgtctcaaacttttaaccatttctttccagaagaaaaa
I L L H L T S L S Q T F N H F F P E E K repeat 1B
tttgaaacattaagggaaaacagttgggtaaaagatccatttgcttttcgacaccctgaa
F E T L R E N S W V K D P F A F R H P E
tcaataattgagctaaacttggtgcctgaagaagagaatgaattattgcagcttagttct
S I I E L N L V P E E E N E L L Q L S S
tcatatacattgaagaatgattatgaaaccttaagtttatcagcattttggatgaaggta
S Y T L K N D Y E T L S L S A F W M K V
aaggaagactttccattgttaagtagaaagagtgtcctgctattgctaccattcacaaca
K E D F P L L S R K S V L L L L P F T T
actagtttgtgtgaactagggttttccatcttaacgcagttaaaaacaaaggaaagaaat
T S L C E L G F S I L T Q L K T K E R N
gggctgaattgtgcagcagttatgcgggtagcattatcttcctgtgttccagactggaat
G L N C A A V M R V A L S S C V P D W N
gaacttatgaacaggcaagcacacccatcatagtaaataaaaatcttacctagcttttgt polyA_signal_sequenc (regulatory), major polyA sight
E L M N R Q A H P S * * I K I L P S F C stop codon
ctttgtatttcttattttgtagtatttttctatgttatatttaaatggtactataatact
gtgatacttttgttatgttttaatttttgttatatttaataaaattattttatgttcatt
gaacaaaaatttaatgaatttctgttagaggccaggaactattctagagacatttgggat
acaaaagtgaacaaaacaggtaattccctagtagagtttatatcctggcaaggagaaatt
gacaataaacctaataaataaggtttataatatttagaagctattaagtgctatggaaag
agtagtaagaaggaaggtcagggaagtactggggaaccaaaccatgaagggttctgtaga
ccattattgggcctctggcttttgtcagtggactagagaacagttgaagggtttaagcga
aggagagaaatgatctgagctaggttttaaaagacactctggtcactattttaaattctt
agggtaagtctgaattaaatgttactttcccctcactgggcatggtggctcagacctgta
atccccgcactttggcaggccatggcagaaggctctgttgagcccaggagttcaagacca
tcctgggcaacatagtgagaccttatttctactaaaaatattttaaaaataagtcaggtg
tggtggtgcacacctatagtcccagctactcaggaggctgtggcaggagggtcgcttgac
ctcaggagtttgaagttgcagtcatctatgattgcaccactgcagtccagcctgagcaac
agagtgaaaccctgtctcgagaaaattaaatgttacttccctaaaaaaaccttttctaac
caccctagggtaaatcctccattattcctttatttctttgttttccttgtagtatataat
ttgtaataattttgattactgattgtcattctgccaccctggagtatataatttttaatt
atctgattactgttattcttccatagtaggggaggtgatatccatttgcctgatacatag
tatgtgttcaatacacatttgctaaagaataaatgaatcaataatacctaacatctctaa
tttgcagtcattcccaagagtaattattaaatatgtggcaaatttctttgcctttttact
tttaaaaatctaattttgacataactgctgtaaccatccagaaacggcattgatgttgct
tcacgttgctgatgcttaagcaatgtatattgtgtaatatacaatgtagtcttcaaacta
atttcaacttctgcctttctgtgtactcccttatcccactgggtgatattatttggcatg
gtcattgtcattaaaatcatacaggatagtaattcctttccatctgctaccatgcctagc
cttatttaatttttcagattttctgttctattgaaggtaattgattttttcttttttttt
atgcttgaaataaagtgttgaaaaacaa] polyA_signal_sequenc (regulatory), major polyA sight [exon2]
- Global Sequence Alignment
There were no orthologs for fish, fruit fly, mouse or chicken as suggested but I found Ascaphus truei, an amphibian with 34% identity.
Content Disclaimer
Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.
- The information displayed on this website is sourced in part or in whole from Wikipedia and has been adapted for the purpose of restating it. We strive to provide accurate and relevant information, however:
- There is no guarantee of absolute accuracy. Wikipedia is an open, collaborative project that can be edited by anyone, so information is subject to change.
- It is not intended to constitute professional advice. The content displayed is for informational and educational purposes only. For important decisions (e.g., medical, legal, or financial), please consult a professional.
- Content copyright. Wikipedia is licensed under the Creative Commons Attribution-ShareAlike License (CC BY-SA). This means that content may be reused with appropriate attribution and shared under a similar license.
- Responsible use. Any risk arising from the use of information from this website is entirely the responsibility of the user.