Big Data and Other Challenges in the Quest for Orthologs

Given the rapid increase of species with a sequenced genome, the need to identify orthologous genes between them has emerged as a central bioinformatics task. Many different methods exist for orthology detection, which makes it difficult to decide which one to choose for a particular application.

Here, we review the latest developments and issues in the orthology field, and summarise the most recent results reported at the third 'Quest for Orthologs' meeting. We focus on community efforts such as the adoption of reference proteomes, standard file formats, and benchmarking. Progress in these areas is good and they are already beneficial to both orthology consumers and providers. However, a major current issue is that the massive increase in complete proteomes poses computational challenges to many of the ortholog database providers, as most orthology inference algorithms scale at least quadratically with the number of proteomes.

The Quest for Orthologs consortium is an open community with a number of working groups that join efforts to enhance various aspects of orthology analysis such as defining standard formats and datasets, documenting community resources, and benchmarking. All such materials are available at

Sonnhammer, E., Gabaldon, T., Sousa da Silva, A.W., Martin, F., Robinson-Rechavi, M., Boeckmann, B., Thomas, E., Dessimoz, C., Vandepoele, K., Van Bel, M., Quest for Orthologs consortium, (2014) Big Data and Other Challenges in the Quest for Orthologs. Bioinformatics 30(21):2993-8.

