User:Madzgopher

Problem Set 1

     0) AI Overview

The FAM200B is a protein gene (Family With Sequence Similarity 200 Member B),when in humans is encoded by c5orf53, is a protein-coding gene located on chromosome 4p15.32. The human variant transcript (NM_001145191.2) encodes a 657–amino acid protein(UniProt reviewed entry P0CF97). Its physiological function remains uncharacterized, expression data confirm that its RNA is present across many human tissues, including brain and thyroid, with moderate levels (RPKM ~5–6). Protein-level immunohistochemistry has been attempted, but the Human Protein Atlas reports insufficient reliability for tissue-level estimation.

Proteomics data confirm that FAM200B protein exists in vivo, meeting criteria for Protein Evidence Level 1 (PE1) in neXtProt/ProteomeXchange. In functional studies, a large-scale RNA interference screen showed that siRNA-mediated knockdown of FAM200B inhibited HIV-1 replication in TZM-bl cells (Brass et al., Science, 2008), suggesting a potential but uncharacterized role in viral host interactions. 6

No FAM200B associations are established in ClinGen or OMIM, and FAM200B has no confirmed clinical phenotype links. It has a known paralog (FAM200A) and multiple orthologs across vertebrates, indicating evolutionary conservation, but its molecular role remains undefined.

The AI wasn't very helpful. The first paragraph was information I had already known and determined the  protein was Level 1 meaning it's real and in humans which I already knew. Lastly it gave me a potential lead about function but didn't give me any solid sources (not sure what Brass 2008 was) I could confirm this information. Overall not very helpful but I am going to keep looking into HIV studies.

  1. Bioinformation: Accession number: >NM_001145191.21, Protein: >NP_001138663.12

Aliases: C4orf54

Homo sapiens FAM200B is a 4287 nucleotides gene located on chromosome 4 at 4p15.321 and encodes a 657 amino acid protein2. The gene has 2 exons, the second one consisting of the base pairs 76 to 4287 of the mRNAcontaining almost the whole gene1. The gene has 23 transcript variants encoding 2 protein isoforms. FAM200B transcript mRNA is found in multiple cell lines, human tissues and cancer samples however with low tissue specificity in all categories; meaning it's not specific to a particular tissue or region.5 The gene has orthologs in mammals, birds, reptiles, amphibians and fish. FAM 200B has over — paralog with the most important being FAM200A with 79.79 % identity. Omin and wikipedia concluded no search results however PubMed found some scientific articles containing “FAM200B”. In a Mendelian randomization study using dopaminergic neuron–specific eQTL data, altered expression of FAM200B was identified as significantly associated with Parkinson’s disease risk, meaning it's a potential gene for disease susceptibility.3 FAM200 B showed low tissue specificity for the brain however is present in the basal ganglia, the part of the brain most associated with Parkinson's disease. Within the brain the gene is 3rd most in the basal ganglia behind the hypothalamus and white matter.5 In another study looking at childhood constipation, decreased expression of FAM200B was observed in patients and identified as a hub gene associated with reduced CD56⁺ NK cell counts (p < 0.05).4 Both articles suggest the gene is present in many parts of the body and responsible for many possible functions.

Repeat 1: Expressed in my needle pairwise alignment for Homo sapiens and Homo sapiens. Looked for internal structure of FAM200B and found 2 repeats; figure 1 and 2.

  1. Sequence Bioinformation

> FAM200B - 4289 nucleotides

[nnagaggcgcagctctgctggagagacctgaggcttgcagcgctgggcccggccgggtcg [exon 1]/[exon 2]

tgcggctgggctgga][ggtcaaggatgcattccaggaaagttctataaaaatataccatc

aactccaacctgatatttgacctagccatggaaaacctcattggctcctcactggatcag

accataaaaatgatccaggaatgaaaatctaagataaattcccaccacttcaagatgttc

ttatacaagtcaacagtctagcagatcagccagtctctttcatcatcaagcagacatact

caaggtaccaacacagcaaaggcacaaaatgtctctatatcagtatttttcatggtatat

gtgagacccattagtgaggcatgaaatcaattctgtgagtcacaatgaacatttgtctaa

aaaacaaaataaaatgtaaaataccaaagtgttttgcatgaatcctttgcctgaattgtg

cacattgcaaggaatgctgaggtcaatttcaatatgcgagaggggtatttgatggaggag

atatgatgaatagaatgtatattggctaagaaatggttttcagtcattgattcctgggtg

tcttgcttgattttcataatagcacggttaatgttctatagtcaagttttatatgattacaat

tacttgcaactagttgtgaagtctgaaatattcgattcaattgcaaactttgagtgtagt

ttttgaaaagatgttatttagactgtatatttttttcttcttttttgagttagtgccaat

Tataacattttaatcaaactggaacaaattgctattaaaatggatcatttctttattaaa   start codon

Y  N  I  L  I  K  L  E  Q  I  A  I  K  M  D  H  F  F  I  K

agaaagaggaatagtgaagtgaaatatacagaagcatgttcaagttcatctgttgaatct

R  K  R  N  S  E  V  K  Y  T  E  A  C  S  S  S  S  V  E  S

ggaattgtgaatagtgacaatattgagaaaaatactgactccaatctgcaaacttcaact

G  I  V  N  S  D  N  I  E  K  N  T  D  S  N  L  Q  T  S  T

tcatttgagccacatttcaaaaagaaaaaagtaagtgcaagacgttataatgaagattac

S  F  E  P  H  F  K  K  K  K  V  S  A  R  R  Y  N  E  D  Y

ttaaaatatggctttatcaaatgtgaaaaaccctttgaaaatgacagacctcagtgtgtt

L  K  Y  G  F  I  K  C  E  K  P  F  E  N  D  R  P  Q  C  V

atttgtaataatattcttgcgaatgaaagcttaaaaccttcgaaattaaaaaggcactta

I  C  N  N  I  L  A  N  E  S  L  K  P  S  K  L  K  R  H  L       repeat 2B

gaaactcagcatgctgaacttattgataagcctcttgaatattttcaaagaaagaaaaaa

E  T  Q  H  A  E  L  I  D  K  P  L  E  Y  F  Q  R  K  K  K

gacataaagttatcaacacaatttcttagttgttctactgctgttagtgagaaagcctta

D  I  K  L  S  T  Q  F  L  S  C  S  T  A  V  S  E  K  A  L

ttatcatcatatttagttgcatatcgtgtggcaaaagagaaaatagctaacacagctgct

L  S  S  Y  L  V  A  Y  R  V  A  K  E  K  I  A  N  T  A  A

gaaaaaattattcttccagcatgtttggatatggtgcgtacaatatttgatgataaatca

E  K  I  I  L  P  A  C  L  D  M  V  R  T  I  F  D  D  K  S

gctgataaattaaaaactatacctaatgataacacagtatctcttcgaatttgtactatt

A  D  K  L  K  T  I  P  N  D  N  T  V  S  L  R  I  C  T  I       repeat 2A

gcagaacatttagaaacaatgcttattactcgtttacagtctggtatagattttgcaatc

A  E  H  L  E  T  M  L  I  T  R  L  Q  S  G  I  D  F  A  I

cagcttgatgaaagcactgatattggaagctgcacaacacttttagtttatgtcagatat

Q  L  D  E  S  T  D  I  G  S  C  T  T  L  L  V  Y  V  R  Y

gcgtggcaagatgattttttggaggattttttgtgttttttaaatttaacctcacaccta

A  W  Q  D  D  F  L  E  D  F  L  C  F  L  N  L  T  S  H  L

agtggattagatatttttacagaattagaaaggcgcatagttggccaatataaattaaac

S  G  L  D  I  F  T  E  L  E  R  R  I  V  G  Q  Y  K  L  N

tggaaaaactgtaaaggaattacaagtgatggcacagcaaccatgactggaaaacatagc

W  K  N  C  K  G  I  T  S  D  G  T  A  T  M  T  G  K  H  S

agagtaattaaaaaattactagaagttactaataatggtgctgtgtggaatcattgtttt

R  V  I  K  K  [L  L  E  V  T  N  N  G  A  V  W  N  H  C  F]    repeat 1A

atacatcgtgaaggtttagcatccagagaaattccacagaatctcatggaggtattgaaa

I  H  R  E  G  L  A  S  R  E  I  P  Q  N  L  M  E  V  L  K

aatgcagtgaaagttgttaattttattaaaggaagctcattgaatagccggcttcttgaa

N  A  V  K  V  V  N  F  I  K  G  S  S  L  N  S  R  L  L  E

acattttgttcagagattggaactaatcatacccacttactatatcataccaaaattcgt

T  F  C  S  E  I  G  T  N  H  T  H  L  L  Y  H  T  K  I  R

tggttgtctcaagggaaaatactaagcagggtttatgagctcaggaatgagattcacttt

W  L  S  Q  G  K  I  L  S  R  V  Y  E  L  R  N  E  I  H  F

tttctcattgaaaaaaaatctcatttggcaagtatttttgaagatgatacttgggtaaca

F  L  I  E  K  K  S  H  L  A  S  I  F  E  D  D  T  W  V  T

aaattggcatatttaactgatatttttagcattcttaatgaactgagtttaaaactacag

K  L  A  Y  L  T  D  I  F  S  I  L  N  E  L  S  L  K  L  Q

gggaaaaacagtgatgtattccaacatgttgaacgtatccagggatttcgaaagacatta

G  K  N  S  D  V  F  Q  H  V  E  R  I  Q  G  F  R  K  T  L

ttgttatggcaagtaagacttaaaagtaatcgtcctagctattacatgtttccaagattt

L  L  W  Q  V  R  L  K  S  N  R  P  S  Y  Y  M  F  P  R  F

ttgcagcatattgaagagaatattattaatgaaaacattttgaaagaaataaaattagag

L  Q  H  I  E  E  N  I  I  N  E  N  I  L  K  E  I  K  L  E

atattgttgcatctcacttctctgtctcaaacttttaaccatttctttccagaagaaaaa

I  L  L  H  L  T  S  L  S  Q  T  F  N  H  F  F  P  E  E  K        repeat 1B

tttgaaacattaagggaaaacagttgggtaaaagatccatttgcttttcgacaccctgaa

F  E  T  L  R  E  N  S  W  V  K  D  P  F  A  F  R  H  P  E

tcaataattgagctaaacttggtgcctgaagaagagaatgaattattgcagcttagttct

S  I  I  E  L  N  L  V  P  E  E  E  N  E  L  L  Q  L  S  S

tcatatacattgaagaatgattatgaaaccttaagtttatcagcattttggatgaaggta

S  Y  T  L  K  N  D  Y  E  T  L  S  L  S  A  F  W  M  K  V

aaggaagactttccattgttaagtagaaagagtgtcctgctattgctaccattcacaaca

K  E  D  F  P  L  L  S  R  K  S  V  L  L  L  L  P  F  T  T

actagtttgtgtgaactagggttttccatcttaacgcagttaaaaacaaaggaaagaaat

T  S  L  C  E  L  G  F  S  I  L  T  Q  L  K  T  K  E  R  N

gggctgaattgtgcagcagttatgcgggtagcattatcttcctgtgttccagactggaat

G  L  N  C  A  A  V  M  R  V  A  L  S  S  C  V  P  D  W  N

gaacttatgaacaggcaagcacacccatcatagtaaataaaaatcttacctagcttttgt polyA_signal_sequenc (regulatory), major polyA sight

E  L  M  N  R  Q  A  H  P  S  *  *  I  K  I  L  P  S  F  C        stop codon

ctttgtatttcttattttgtagtatttttctatgttatatttaaatggtactataatact

gtgatacttttgttatgttttaatttttgttatatttaataaaattattttatgttcatt

gaacaaaaatttaatgaatttctgttagaggccaggaactattctagagacatttgggat

acaaaagtgaacaaaacaggtaattccctagtagagtttatatcctggcaaggagaaatt

gacaataaacctaataaataaggtttataatatttagaagctattaagtgctatggaaag

agtagtaagaaggaaggtcagggaagtactggggaaccaaaccatgaagggttctgtaga

ccattattgggcctctggcttttgtcagtggactagagaacagttgaagggtttaagcga

aggagagaaatgatctgagctaggttttaaaagacactctggtcactattttaaattctt

agggtaagtctgaattaaatgttactttcccctcactgggcatggtggctcagacctgta

atccccgcactttggcaggccatggcagaaggctctgttgagcccaggagttcaagacca

tcctgggcaacatagtgagaccttatttctactaaaaatattttaaaaataagtcaggtg

tggtggtgcacacctatagtcccagctactcaggaggctgtggcaggagggtcgcttgac

ctcaggagtttgaagttgcagtcatctatgattgcaccactgcagtccagcctgagcaac

agagtgaaaccctgtctcgagaaaattaaatgttacttccctaaaaaaaccttttctaac

caccctagggtaaatcctccattattcctttatttctttgttttccttgtagtatataat

ttgtaataattttgattactgattgtcattctgccaccctggagtatataatttttaatt

atctgattactgttattcttccatagtaggggaggtgatatccatttgcctgatacatag

tatgtgttcaatacacatttgctaaagaataaatgaatcaataatacctaacatctctaa

tttgcagtcattcccaagagtaattattaaatatgtggcaaatttctttgcctttttact

tttaaaaatctaattttgacataactgctgtaaccatccagaaacggcattgatgttgct

tcacgttgctgatgcttaagcaatgtatattgtgtaatatacaatgtagtcttcaaacta

atttcaacttctgcctttctgtgtactcccttatcccactgggtgatattatttggcatg

gtcattgtcattaaaatcatacaggatagtaattcctttccatctgctaccatgcctagc

cttatttaatttttcagattttctgttctattgaaggtaattgattttttcttttttttt

atgcttgaaataaagtgttgaaaaacaa]    polyA_signal_sequenc (regulatory), major polyA sight [exon2]


  1. Global Sequence Alignment

There were no orthologs for fish, fruit fly, mouse or chicken as suggested but I found Ascaphus truei, an amphibian with 34% identity.

Content Disclaimer

Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.

  1. The information displayed on this website is sourced in part or in whole from Wikipedia and has been adapted for the purpose of restating it. We strive to provide accurate and relevant information, however:
  2. There is no guarantee of absolute accuracy. Wikipedia is an open, collaborative project that can be edited by anyone, so information is subject to change.
  3. It is not intended to constitute professional advice. The content displayed is for informational and educational purposes only. For important decisions (e.g., medical, legal, or financial), please consult a professional.
  4. Content copyright. Wikipedia is licensed under the Creative Commons Attribution-ShareAlike License (CC BY-SA). This means that content may be reused with appropriate attribution and shared under a similar license.
  5. Responsible use. Any risk arising from the use of information from this website is entirely the responsibility of the user.