Note

Bioinformatics is the combination of biology and informatics (not to be confused with computer science). It’s used for “in silico” analyses of biological data (from “in vivo” or “in vitro” sources).

Bioinformatics could be defined as the science of information and information flow in biological systems, especially of the use of computational methods in genetics and genomics.

As biology itself is extremely complex, we’ll see a simplified view.

Words ending with “ome” mean “the complete set of”, while words ending with “omics” mean “the study of”.

Cells

Note

A cell is the smallest structural and functional unit of an organism, which is typically microscopic and consists of cytoplasm and a nucleus enclosed in a membrane.

There’s many types of cells, such as animal cell, plant cell, bacteria cell, etc.

center

Cells are made out of molecules and proteins.

Cellular signalling and metabolic processes

Cells have surface receptors that allows signals to be recovered from outside itself. center

Cells may also execute some metabolic processes, which can be:

  • Catabolic: break down components
  • Anabolic: synthesise new components

Molecules

Molecules could be sources of energy, “signaling molecules” for signal transmission or building blocks of macromolecules.

Proteins

Proteins are macromolecules, and they can be considered the building blocks of the cell. They are used for biological processes and signal transduction.

They can often join in protein complex, or have otherwise physical “protein-protein interactions”.

DNA and RNA

DNA (DeoxyRibonucleic Acid) is a macromolecule that contains genes, regulatory elements and “junk DNA”. It has a double helix shape and it’s used for storage and reproduction of information.

center

RNA (RiboNucleic Acid) is a macromolecule that has a key role in transformation of genetic information to function. Differently from DNA, it’s single-stranded.

DNA and RNA

Note

Chemically speaking, DNA and RNA are a polymer made out of nucleotides. Nucleotides are made of different components: a sugar group, a phosphate group and a base.

There are four different bases for DNA: Adenine, Thymine, Guanine and Cythosine. As for RNA there’s the dame bases, except for Thymine which is replaced by Uracil.

center

It’s important to know that the backbone is formed by covalent bonds.

In the DNA double helix structure it’s important to know that A pairs with T and C pairs with G.

Chromosomes

DNA is usually organised in chromosomes (there’s some exceptions like prokaryotes).

The diploid number (number of chromosomes) and genome sizes differ widely between species.

Human chromosomes come in pair, and can be either XX or XY.

Reverse and complement operations

We define the complement of a DNA strand as the same strand where we replaced A by T, C by G, G by C and T by A.

The reverse is just the same strand in reversed order.

The reverse complement of a DNA strand gives the opposite DNA strand, this is the principle behind DNA replication, which is done by reconstructing the opposite strand for each of the two separated original DNA strand.

center

DNA stores the information that determines a protein’s structure. DNA is later transcribed into RNA, which mediates the transformation of genetic information into functional molecules. RNA is later translated into a protein, which exerts a biological function.

Proteins

Note

Proteins are involved in the most of the tasks essential for life:

  • Structural proteins
  • Receptor and channel proteins
  • Signalling proteins
  • Enzymes
  • Transcription factors

Transcription process

In the transcription process, one strand of DNA is copied into a reverse complementary RNA molecule, this process is executed by a RNA polymerase. In prokaryotes cells it’s transcribed directly into mRNA, while in eukaryotes cells, it’s transcribed in pre-mRNA.

In eukaryotes cells, transcribed DNA is often spliced before it’s exported from the nucleus to the cytoplasm. Alternative splices of transcripts allow the production of different gene products from the same gene.

Translation process

mRNA transcripts are later translated into proteins by the ribosome, which encodes amino acids into “codons”, which are nucleotide triplets.

The translation starts at AUG and ends at the first UAA, UAG or UGA.

Each codon is recognised by a tRNA which carries the correct amino acid.

Viruses

Viruses are organisms that exploit a host cell for replication. It attaches to a cell and injects some viral DNA, which is later replicated and released outside.

DNA Sequencing

Note

There are various techniques to sequence the DNA.

The sequencing by synthesis technique starts from a single stranded DNA, an then reconstructs the reverse strand and checks which bases you have to add to reconstruct it. This was done by marking nucleotides by using radiation. This is now done by using different fluorescent dyes to read nucleotides.

Next Generation Sequencing

Next Generation Sequencing (NGS) refers to one of a number of technologies that enable a massive parallelisation of DNA sequencing.

One of the best NGS is Illumina sequencing. This approach uses 4 steps:

  1. DNA and library preparation: make random DNA fragments and append adapters
  2. Chip/Flowcell preparation: Attach fragments to surface and amplify
  3. Sequence: massively parallel DNA sequencing
  4. Analyze