SEMARANG – The use of bioinformatics is becoming increasingly important in managing and interpreting large-scale biological data, particularly in the fields of medicine and genomic research. This was one of the key focuses of the “Ensembl Training Events” series organized by Faculty of Medicine Universitas Diponegoro (FK UNDIP), featuring Aleena Mushtaq, Senior Ensembl Outreach Officer, as trainer. The event took place on September 4, 5, and 7, 2026, at the Senate Meeting Room, 2nd floor of the Dr. Soewondo Building, FK UNDIP.
The training series included “The Ensembl Data Platform for Clinicians,” “The Ensembl Data Platform Training,” and “The Ensembl Train the Trainer Workshop.” The sessions were designed to introduce participants to various Ensembl resources while enhancing their understanding of how to access, explore, and analyze genomic data for both clinical and research purposes within FK UNDIP.
During “The Ensembl Data Platform for Clinicians” session, participants were introduced to Ensembl, a genomic resource that provides genome data, gene annotations, genetic variation, regulatory elements, and comparative genomics information. Through presentations and hands-on practice, participants learned how genomic data can be used to derive more meaningful biological insights in a medical context.

During the introductory session, Aleena Mushtaq introduced the basic concepts of bioinformatics as a field that integrates biology and computer science to manage, analyze, and interpret large volumes of biological data. In medicine, bioinformatics can be used to analyze genetic profiles, help assess disease risks, support diagnosis, and study patients’ responses to medications.
The workshop also explored the role of Ensembl in transforming raw sequence data into more interpretable genomic information through genomic annotation. Through this process, features such as genes, variants, and regulatory elements can be mapped onto genome sequences, providing biological context for genomic coordinates.
Participants were then introduced to various genome assemblies, including GRCh37, GRCh38, telomere-to-telomere (T2T), and the Human Pan-Genome Reference Consortium (HPRC). Selecting the appropriate genome assembly is an important part of genomic analysis, as the genomic coordinates of a gene or variant may differ between assemblies.
During hands-on practice with the Ensembl Genome Browser, participants learned how to select species and genome assemblies, search for specific genes, and examine gene positions within the genomic context. One example used during the session was the BRCA2 gene in the human GRCh38 genome. Participants also explored transcripts, exons, introns, regulatory elements, and various types of variation that can be displayed through the Genome Browser.
In addition to BRCA2, participants practiced using the MCM6 gene to understand gene orientation and identify protein-coding genes located downstream. The exercise was conducted interactively using the Slido platform, allowing participants to immediately test their understanding of the material presented.
The session then continued with a discussion of genetic variation and variant consequences. Participants were introduced to various types of variants, including single nucleotide variants (SNVs), insertions and deletions (indels), and structural variants. The workshop also explained how the position of a variant can determine its molecular consequence, including synonymous, missense, splice, UTR, regulatory, and intergenic variants.

In the following session, participants were introduced to the Ensembl Variant Effect Predictor (VEP), a tool used to annotate variants and predict the functional consequences of variants identified through sequencing analysis. Participants learned how data in VCF, HGVS, and other formats can be submitted to VEP to obtain information on affected genes and transcripts, variant consequences, population frequencies, protein information, clinical significance, and relevant publications.
VEP also provides several additional analysis options, including computational predictions such as SIFT, PolyPhen, AlphaMissense, REVEL, and CADD. The workshop emphasized that these scores are computational predictions rather than experimental evidence. Therefore, prediction results should be considered alongside other available evidence and, when necessary, validated experimentally.
During “The Ensembl Data Platform Training,” participants gained further insights into using the Ensembl platform to explore and manage genomic information. For large-scale analyses, participants were also introduced to accessing Ensembl through the command line and using various available plugins. This approach provides greater flexibility than a web-based interface, particularly when working with large-scale genomic datasets.
The workshop also highlighted the importance of maintaining genome assembly consistency in research. Researchers need to ensure that variant data and other datasets used in a project refer to the same genome assembly. Participants were also reminded to specify the Ensembl version used when publishing their findings, enabling other researchers to reproduce the analysis using the same data version.
Ensembl currently uses a release system to maintain updated genomic data. The integrated release provides a more stable version for long-term research and reproducibility. During the workshop, participants worked with Ensembl integrated release 2026-07.
Meanwhile, “The Ensembl Train the Trainer Workshop” provided participants with an opportunity to deepen their understanding while gaining the skills needed to share knowledge and introduce the Ensembl platform to other users. The session formed part of efforts to expand the use of genomic resources by strengthening participants’ capacity at FK UNDIP.
Through the Ensembl Training Events, participants gained an understanding of Ensembl as a bioinformatics resource for connecting genomic data with biological information, including genes, genetic variation, regulation, diseases, and drug responses. The sessions also provided practical foundations for using the genomic browser and VEP to support genomic analyses, particularly in research and clinical applications at FK UNDIP.
Ensembl provides open access to a wide range of genomic data, along with tools for exploration, analysis, and programmatic access. As the volume of genomic data continues to grow, proficiency in bioinformatics resources such as Ensembl can contribute to strengthening data-driven medical research at FK UNDIP. (Humas FK UNDIP)




