Episode 211: Coefficient of Inbreeding

Adam Boyko, PhD explains everything you need to know about COI

Good Dog is on a mission to educate the public, support dog breeders, and promote canine health so we can give our dogs the world they deserve.

Good Dog is on a mission to educate the public, support dog breeders, and promote canine health so we can give our dogs the world they deserve.

Good Dog is on a mission to educate the public, support dog breeders, and promote canine health so we can give our dogs the world they deserve.

Advances in genomics have made accurate COI measurements easier than ever, but understanding how to interpret these values and why they often differ substantially from pedigree-based COI estimates can be tricky. Adam Boyko, PhD does a deep dive into genetics to help you learn all about COI – what it measures, how it’s calculated, and why it’s important.

Watch the video version of this presentation here.

Transcript

Nicole Engelman  00:04

Welcome to the Good Dog Pod. Join us every other Wednesday when we discuss all things dogs, from health and veterinary care to training and behavior science, as well as the ins and outs of Good Dog, and how our platform can help you successfully run your breeding program. Follow us and join Good Dog's mission to build a better world for our dogs and the people who love them. Thank you all so much for being here for our presentation on the Coefficient of Inbreeding (part one) with our guest, Dr. Adam Boyko. This is part one of a two-part series, which we'll be continuing later on this month, so just a little bit about what we're going to cover today. We're going to talk about how advances in genomics have made accurate COI measurements easier than ever, but understanding how to interpret these values and why they often differ substantially from pedigree-based coefficient of inbreeding estimates can be pretty tricky. So today Dr. Boyko is going to help us dive deeper into genetics to help everyone here learn all about the coefficient of inbreeding, what it measures, how it's calculated, and why it's important. And just judging by the amount of people we have here today, it sounds like this is a topic people are super interested in learning more about. So we're excited that this is a two-part series for everyone. We want to, of course, thank our friends at Purina for helping us bring this webinar to you and bring it to life, and, of course, to all of you for sharing with us throughout the past few months what you want to learn more about in 2026. This is something that we haven't really covered in depth so far, and it's a reflection of, again, all of your suggestions, so keep sending those topic suggestions our way. As I'm sure many of you know, we've hosted a ton of webinars this year already, so I just want to redirect everyone to where you can find all of those recordings, plus ones from years past, as well, because we have a really great library built up with really, really great speakers covering really helpful topics, so that is always there for breeders whenever you need it. And as usual, during the Q&A section of this presentation, we're going to prioritize some previously submitted questions from our community first. Great, I think those are the big housekeeping things, but I'm guessing maybe we have some breeders that are new to Good Dog, so I just want to share a little bit about who we are. Good Dog is on a mission to build a better world for our dogs and the people who love them by advocating for breeders like yourselves, educating the public, and promoting canine health and responsible dog ownership. We are an online community created just for breeders like yourselves to connect with Good Dog buyers from all across the country and find forever homes for your puppies. We help readers run all aspects of their program, from getting your new litter of puppies listed, connecting with Good Dog buyers, getting to know them, and receiving payments securely, and all of those smaller things that happen in between your puppies being born and going to their new homes. We're really here to support you every step of the way through that journey. We also have a number of educational resources—that's something really important to us—and health-related discounts to help your programs thrive in other ways as well. So, if you're not yet a member of our community, we'd love to have you learn more about our mission and apply to join at gooddog.com/join. And then the very last thing for me, because I want to make sure we have plenty of time for this presentation. I want to talk a little bit more about Dr. Boyko, our speaker today, and his amazing background in canine health. Dr. Boyko is an associate professor in the Department of Biomedical and Translational Sciences at the Cornell University College of Veterinary Medicine in Ithaca, New York. (I went to Syracuse, so I'm a big fan of Ithaca. I think it's a beautiful city.) Dr. Boyko's research focuses on canine genomics, including the evolution and population history of dogs, the genetics of canine traits and disease, and the role of genetic diversity and inbreeding depression on population health. He's also the co-founder of Embark Veterinary, a dog genetic testing company, which I'm sure our community is very familiar with, and with that, Dr. Boyko, I'll pass things over to you to start our presentation.


Adam Boyko, PhD  04:08

Great! Thanks, Nicole. Thank you all for joining today's webinar on coefficient of inbreeding. It's a subject I spend a lot of time thinking about, particularly in the context of dog breeding. Now, I'm a geneticist by trade. I'm not a dog breeder. So some of this is going to be a little technical as we dig into the details, but I'll do my best to explain it all and give everyone time for questions at the end. Now, coefficient of inbreeding is usually referred to as COI, is familiar to many, probably most breeders, but it's actually a pretty complicated and confusing metric, because it's something that can be measured a few different ways, and those ways don't always agree with each other when we look at them in real animals. So at its core, the coefficient of inbreeding is a way to measure inbreeding in an animal, and we all kind of—dog breeders, dog owners, even cat lovers—everyone has an idea of what inbreeding is. Inbreeding is the result of crossing close relatives, like first cousins or brothers or sisters, or maybe even parents and offspring, and when we think about it in humans, we often joke about, you know, poor back-wood folks, like the characters in Deliverance, but we also, on the other end of the spectrum, associate it with things like the incestuous royal families of Europe, and so, like the Habsburgs that used marriages between close relatives to consolidate and control the family's wealth and power. As inbreeding built up over the generations, so did the prevalence of inherited disorders like hemophilia, as well as other deleterious traits. You can see here the pronounced Hapsburg Jaw in King Charles the Second of Spain. So King Charles was known as El Hechizado, or the Bewitched. His underbite was so severe that he couldn't properly chew food, and he didn't learn how to speak until he was four years old. And remember, this is a portrait—not a picture—of Charles the Second, so this is an image drawn by an artist doing his very best to paint his king in a flattering light, and so this is what inbreeding did to him in the Habsburg line. He, despite two marriages, was unable to produce an heir, and he sickened and died at the age of 38. So clearly inbreeding can be bad, at least in humans, but there are some misconceptions about why inbreeding is bad, and whether it's always bad, and how to interpret measures of inbreeding like heterozygosity and COI. 


So we're going to talk about these subjects today, and in the companion seminar later, starting with the genetics of why inbreeding is bad. So if you ask a geneticist why inbreeding is bad, they're going to say that all populations, even genetically healthy ones, contain rare recessive deleterious genetic variants. Now, rare is pretty straightforward. We're talking about genetic variants that are mostly uncommon in the population. Genetically healthy populations have lots of genetic variants, and most of these are rare, meaning only a small fraction, often 1% or less, of the individuals carry the rare variant, which we also call an allele. If we look across dogs as a species, there's a lot of genetic diversity, tens of millions of known variants across their two and a half billion base pair genome, and a large fraction of them are rare, so this is a meta-analysis of over 15,000 dogs that was just published last month in the bio archive, but we see that over 30% of the genetic variants that they have in their data set are at less than 5% frequency, so here we're saying it's a minor allele frequency, so when you've got a genetic variant, you've got two different alleles that a chromosome can carry, and the more common one we call the major allele, and that's going to be at 50% frequency or above, and the less common one we call the minor allele, so that's at 50% frequency or below. Now, recessive is a little bit trickier, but it just means that in order for an allele to have an effect on the dog's phenotype—which is the way the dog looks or acts, or things like its susceptibility to certain diseases or its metabolic rate—in order for a recessive allele to affect phenotype, all of the copies of the gene that the dog has have to have the recessive allele. So, remember, for over 95% of the genome, a dog is going to inherit two copies of each chromosome—one from its mother, one from its father—with the exception mainly being the sex chromosome, where in males, where males only have one copy of the X chromosome that they inherited from their mother, and then they have the Y chromosome that they inherited from their father, but otherwise we have two copies of all of the genes that we have in our genome. Finally, there's deleterious, so a deleterious genetic variant is one that can be harmful either by causing a disease like hemophilia or simply contributing to disease susceptibility or by contributing to a negative trait like the Hapsburg Jaw. So for a deleterious recessive allele to be harmful, an individual usually has to inherit two copies of the allele, one from its mother and one from its father, except for the exceptions like recessive males. Now, when you have a population that contains lots of rare recessive deleterious genetic variants, when you practice random mating or outcrossing, very few, if any, animals wind up expressing these deleterious phenotypes. It can happen, but it's rare. For example, in my area, white-tailed deer are at high population densities throughout upstate New York, and they harbor rare recessive genetic variants that can cause different phenotypes, like albinism. Actually, here—not to get too far in the details, but this isn't technically albinism, because the deer still have brown eyes—but I'm just going to shorthand call it albinism to work for kind of the example that we're doing now. I've probably seen over 1,000 white-tailed deer in and around Cornell since I started teaching here in 2011 and I've yet to personally see an albino deer here, even though the carrier frequency for albinism is around 1%. So, what this means is there's a 1% chance that any normal looking deer I see is carrying one copy of the albinism gene, and if these deer happen to mate with other carriers of that albinism gene, their fawns have a one in four chance of inheriting two copies of the rare albino variant, and then therefore they would be albino. With a 1% carrier frequency and random mating, only 1% of 1% of the pairings will be carrier/carrier crosses. So, the frequency of deer with the albino phenotype at birth is only going to be about one in 40,000. Almost the only way to encounter an albino deer in the flesh around here is to head down the road to the old Seneca Army Depot in Romulus. So, this was an army base that was enclosed by a fence during World War Two, trapping in a population of a few hundred deer that the army managed through controlled hunting. After a few years, one or two of the white fur deer popped up in the herd, and the depot commander ordered the hunters not to kill any of the white fur deer, and over the decades only brown deer were hunted, and the white deer started popping up more and more frequently, first slowly, then a little bit more quickly, and now they make up about 40% of the herd. So, if you select for that variant, you can get it, but generally it's being kept at a low frequency, 1% or less, and so you almost never under outcrossing see affected individuals being produced. Now, inbreeding changes the math. If, instead of random mating, close inbreeding was occurring, perhaps sibling mating, then a carrier doesn't have a 1% chance of crossing with another carrier, but rather a 25% chance, because that brother or sister of a carrier is also more likely to have inherited a copy of the recessive albino allele. So, if the deer around Cornell were practicing inbreeding like this, the rate of white fur deer in town would go up 25 fold to about one in 1,000, or one in 2,000. So, the genetic explanation is that with lots of rare recessive deleterious alleles, every individual carries several different deleterious recessive alleles, but the odds of mating with another individual that carries one of the same deleterious recessive alleles is quite low if that individual is unrelated. 


So every person carries several highly deleterious recessive alleles. In fact, most of us carry one or two recessive alleles that are so harmful that it would lead to complete sterility or death if they were present in two copies. If individuals find an unrelated mate, they'd have to be extraordinarily unlucky for them to carry one of the same rare deleterious recessive alleles, so carrier/carrier crosses are unlikely to begin with, and then the probability of producing affected offspring is even less likely, because then it's just a one in four chance for each pairing. The main exceptions are if you have recessive disease alleles on the X chromosome, because you only need one copy for males to express the phenotype, and also what geneticists call compound heterozygotes, and so these are individuals that inherit a recessive disease allele in a gene from one parent, and then they inherited different recessive disease allele for the same gene from the other parent, so the parents are unrelated, but they happen to have the same gene broken by different mutations, and the unlikely offspring inherited broken copies from both parents. But other than these exceptions, you mostly only get recessive diseases and deleterious recessive phenotypes when you have inbreeding, because inbreeding is what brings the same rare deleterious disease recessive alleles together. Now, of course, this doesn't really tell us why pretty much every animal population, even genetically healthy ones have so many rare recessive deleterious genetic variants, and for the purposes of this webinar, I'm not going to get into the nitty gritty details of it, but essentially it's that the genome is pretty big, so even though DNA replication is pretty good, there's lots of new mutations every single generation, and these mutations are more likely to be deleterious than beneficial, and these deleterious mutations can be recessive or dominant, or what we call additive, where heterozygous individuals have a phenotype in between that of the two homozygous classes, but lots of them are recessive because they break a gene or make it less efficient, but having a second working copy of the gene is enough to ensure that the gene functions properly in the organism. And importantly, deleterious mutations that are additive or dominant tend to get weeded out pretty quickly by selection, but deleterious recessive alleles, like the white fur gene that we have in upstate New York, persists when it's rare; it's hidden in the population for a long period of time. So, there's going to be a lot of these kinds of mutations in the population. So, this is just one example of many rare recessive deleterious alleles that are almost certainly lurking around in any population and ready to cause harmful phenotypes if inbreeding exposes them. 


So now let's talk about measuring inbreeding. Traditionally, inbreeding coefficients were measured using pedigrees. So here's what a pedigree looks like when parents are half siblings and the individual only has three unique grandparents. Here's a pedigree where two of the grandparents are half siblings, and an individual only has seven great grandparents. These are both examples of inbreeding, but to different degrees. And so the way traditionally we would measure the coefficient of inbreeding from a pedigree was to look for pedigree loops. So, in the first case, we have a pedigree loop here that contains three ancestors, so the expected inbreeding of the offspring in this loop is one half to the third power. Half of the genome of dog 2 came from dog 1, half of the genome of dog 3 came from dog 1, and there's a 50% chance if both dog 2 and dog 3 inherited the same part of the genome from dog 1 that they inherited the same copy of the DNA because dog 1 itself has two copies of all of its chromosomes, except for the sex chromosomes. In the second case, the loop is five ancestors long, so it's expected to lead to about 3% inbreeding in the offspring. 


Real dog pedigrees are complex, they often contain many loops, but they're specialized software that can be used. It's readily available to compute the inbreeding coefficients from even the most complicated pedigrees, so we don't have to do it by hand. In this pedigree of Nova Scotia Duck Tolling Retrievers, breeders have tried to avoid close matings. This individual here has two parents and four grandparents, but there are some shared ancestors further back, like Baron Fields Rocket Torpedo, so he's part of this pedigree loop containing six ancestors, so we know that the inbreeding coefficient of this individual is one half to the sixth power, at least. Northeast Windsong is also a shared ancestor, so this is also going to contribute to the inbreeding coefficient of the individual. This is also an inbreeding loop containing six ancestors, so essentially the software just goes through, looks for all of the different pedigree loops, and then sums them up to calculate the expected inbreeding coefficient. People still use pedigree software to do this, but most people have switched to using genomics to calculate coefficient of inbreeding, because it offers several advantages. So, let's look at how a geneticist thinks about shared ancestry in a genome. So we can plot out your dog's genome, so it would be 38 pairs of chromosomes, plus either a pair of X chromosomes or X and Y. So here we plotted out some of the chromosomes, and we've done what we call painting them, so we're painting them based on the ancestor that donated that section of DNA to them. So with mom and dad, it's simple, because one copy of each gene of each chromosome came from each parent. So, in this case, we're assuming the dog is female, because we have two X chromosomes, and for each pair of chromosomes, one individual did one copy of DNA, the other parent contributed a different DNA copy. Now, of course, if we look at the nucleotide level, if we're looking at the A's, C's, G's, and T's that make up the DNA code, even different DNA copies will be 99.9% identical. So, we're not saying that mom and dad's DNA copies are completely different, we're just saying they aren't completely identical. With grandparents, it gets slightly more complicated, so half of your dog's mother's genome came from her maternal grandmother and half from her maternal grandfather, but these genomes went through recombination during meiosis. So, the maternal genome of your dog inherited is a mosaic of each of those grandparents, and because of this randomness, your dog is only approximately 1/4 related to each grandparent. It will actually be a bit more or a bit less related to some grandparents by chance. So, nevertheless, if mom and dad were not related, the two copies of DNA your dog inherited will be different throughout the genome. But what if your dog's parents weren't unrelated? What if mom and dad were half siblings that had the same father? In this case, your dog would have the same maternal and paternal grandfather. It's the same individual. And now, if we look at the genome, large swaths of the genome are inherited from the same ancestor and are painted the same color. Approximately half of the maternal genome and half of the paternal genome came from the grandfather. So, approximately 1/4 of the genome has two copies of his DNA. Now, grandpa, of course, had two copies of the genome himself, so it's possible he passed along different copies of the genome to each parent, but there's a 50% chance for each section of DNA passed down from him he actually passed down the same copy. In this case, mom and dad will have identical DNA copies because it was passed down from the same DNA molecule from the same ancestor. 


Now, what does this look like when we look at real data, like the type that DNA testing companies generate with their genotyping arrays and use to calculate inbreeding? Real genetic data is a string of genotypes that algorithms can take as input and try to infer which alleles belong to one copy of the dog's genome inherited from one of its parents, and which alleles belong to the other copy of the dog's genome inherited from the other parent. When both parents transmit the same copy of DNA, because they inherited the same copy of DNA from a shared ancestor, it looks like a long run of homozygous genotypes. We call these regions identity by descent, and the genotypes in these runs are homozygous, because the alleles come from the same ancestral DNA molecule in the recent ancestor. This isn't to be confused with homozygous genotypes in general, which can occur anywhere in the genome. Homozygous genotypes themselves are not clustered in long runs and do not indicate inbreeding, so while overall genome wide homozygosity tends to increase as individuals become more and more inbred and overall genome wide heterozygosity tends to decrease, other factors could also influence heterozygosity and homozygosity, and so they're not equivalent measures to CLI. With identity by state, markers or homozygous simply because by chance, as alleles segregate in a population, they paired up together—not because they descended from the same DNA molecule in the common ancestor. A few adjacent markers can be identical by state, but this doesn't really tell us anything about the untyped markers in that genomic region, and whether to expect them to be homozygous as well. In contrast, long runs of homozygosity indicate identity by descent. We know that every locus in this region of the genome, even for untyped markers, will be homozygous because the whole genomic region was inherited from the same ancestral DNA molecule. Importantly, this means that if the DNA molecule contained an unknown rare recessive deleterious allele, the dog that has this run of homozygosity will be homozygous for that recessive allele, even though we didn't directly genotype the unknown variant. These identical by descent regions of the genome are the hallmarks of inbreeding and occur not only when matings of close relatives occur, but also from more distant relationships, especially when there are few breeding individuals in a population. This is because to avoid inbreeding, an individual needs to have two parents, four grandparents, eight great-grandparents, and so on. So, by 10 generations back, an individual has 1,024 different ancestors, and by 20 generations back, it'll have over a million ancestors. This is, of course, impossible in a small population with only a few hundred breeding individuals. The smaller the population, the sooner individuals in the population will have ancestors in common when they go back in their family tree, and the more inbreeding they'll have overall in their genomes, even without crossing really closely related relatives. 


These IBD segments that are the hallmarks of inbreeding are also the measure of inbreeding. By calculating the cumulative proportion of the genome that is identical by descent, we can determine the coefficient of inbreeding, or COI, that population geneticists traditionally refer to with the small letter f. The way to do this is to define some threshold for what we consider a long run of homozygosity. Stretches of homozygous genotypes beyond this threshold are considered identical by descent and contribute to inbreeding, while stretches below this threshold are considered only identical by state, and they don't contribute to COI. How long should this threshold be? There's no firm answer. It depends on your data set. It depends on your population. Many population studies use one megabase as the threshold, so 1 million base pairs that are identical, but this can miss out on some shorter IBD segments that can contribute to inbreeding. 


Embark has a dense enough genotyping platform that they actually use a smaller 500 kb unit as the threshold when making COI calculations. Once we've chosen this threshold, we can find all the runs of homozygosity larger than that threshold, sum up their lengths, and divide it by the total length of the genome to compute the COI. So, here's an example of an analysis that was done for a Labrador at the Cornell Biobank. This Labrador has an inbreeding coefficient of 17% which is typical inbreeding coefficient for Labs, and you can see how there is some level of inbreeding, some stretches on every single chromosome has some amount of IBD, but it varies substantially between chromosomes. The smaller the threshold, the greater the ability to detect shared ancestors further back in the family tree, but the trade off is the risk of not having enough variable sites in the region to rule out simple identity by state. During meiosis, one recombination event occurs for every about 100 million base pairs of the genome, so there's approximately a 2% chance of a recombination event occurring in any one megabase region of the genome each generation, so there's a 1% chance of recombination on the maternal copy, and there's a 1% chance of recombination on the paternal copy, so using a threshold of a half to one megabase means we still have good power to detect IBD segments inherited from ancestors 20 to 30 generations back, because it's unlikely recombination has broken up these segments below our detection threshold. Especially with genotyping arrays, we'll want to also define the minimum number of markers in a region before we conclude a run of homozygosity as an indicator of identity by descent and contributes to the coefficient of inbreeding. This is because some portions of the genome might not be well covered by an array and will just have a few markers. They might happen to be identical by state by chance. There's other complexities to consider as well, such as the possibility of genotyping error breaking up a parent run of homozygosity. Lots of little details you don't have to worry about, but you should be aware of, because it can make the comparisons across different companies or different labs be different. 


Calculation of these long runs of homozygosity are actually done in genomic studies for two reasons. The first, and the one we're most interested in today, is to compute COI, but the second is for trying to identify recessive alleles through homozygosity mapping, so I'm just going to take you through an example of homozygosity mapping, so you can understand this technique, because not everybody's familiar with it. So, homozygosity mapping requires both a pedigree and genetic data. A homozygous recessive phenotype, let's say, has appeared in some family, and some individuals within that family exhibit the phenotype. So we know the dogs with the phenotype are homozygous for the unknown recessive variant, and we also know their parents have to be carriers, otherwise they wouldn't have inherited two copies of the recessive gene. So we'll genetically analyze each of the individuals in the pedigree and identify the long runs of homozygosity contained in each of their genomes, and these runs can be used to identify the chromosomal region or regions that must contain the recessive variant. So, here we've lined up the genomes of the sire and the dam, who are obligate carriers, and three of their offspring. The two on the end show the recessive phenotype, so they have to be homozygous for the recessive allele, and the one in the middle doesn't show the phenotype. So, we've colored the runs of homozygosity red, and now the question is: Can you identify the run of homozygosity that likely contains the recessive allele? This portion of the third chromosome is the region of interest for finding the recessive variant, and this is because it's in a long run of homozygosity in both of the affected pumps. And importantly, it's not in a run of homozygosity in either the parents, because the parents are obligate carriers of the variant, or heterozygous for the variant, they should not have a run of homozygosity in the region where the variant occurs. Now, in this example, the unaffected litter mate also doesn't have a long run of homozygosity in this region, but, of course, it's possible the unaffected litter mate does inherit a long homozygosity run in this region in case both parents are heterozygous in the same way, so both parents have one identical copy of the gene with a genetic defect and one identical copy of the gene without the genetic defect. Then, if the littermate inherits two identical copies of the gene without the genetic defect, it'll have a long run of homozygosity, but it won't be phenotypically affected, because it won't have any copies of the recessive allele. If we look further down on chromosome five, we see another region of the genome where a long run of homozygosity is shared by the two affected pups, but in this case, we can rule out this as being a likely region of the genome containing the recessive allele, because one of the parents also has a long run of homozygosity, so because each parent is an obligate carrier, neither should have a run of homozygosity in the genetic region where the variant occurs. 


Okay, so what are some advantages to using genetic COI instead of pedigree-based COI? So, there's a lot of advantages. One of them is you should probably be doing genetic testing anyway. We have lots of known diseases in most breeds, you probably want to be testing to make sure you don't have carriers being crossed with carriers, and it's easiest to do panel testing anyway. It's readily available, and so that's going to generate the data that you need in order to calculate the COI. But beyond that, many populations don't have a pedigree. Now, this isn't something dog breeders typically have to deal with. Dog breeders have really extensive pedigrees that are really great for lots of things, but if you're doing conservation projects—when I work with biologists, they're not going to have a pedigree. They need to find a genomic way to calculate inbreeding for their populations. The pedigree could be wrong, cases of non-paternity happen, so you can't fool the DNA molecules, but you can certainly fool the pedigree software. Pedigrees also only capture identity by descent, going back as far as the pedigree is recorded, so often it's only five to 10 generations, right? But as we've seen, our ability with genotyping arrays, we can go back much further and capture nearly all of the identity by descent caused by shared ancestors 10, 20, 30 generations back. And this is really important, particularly for dog breeds. So, let's look back at the Toller pedigree that we were examining before. If you zoom out and you take a look at the whole pedigree, like every other purebred dog pedigree, you're going to get this—what we call a pedigree collapse. There's this bottleneck that happens at breed formation, and so you've got the same DNA being shuffled around because there's only a few different ancestral dogs when the breed was founded. So all of these dogs are going to have really high pedigree COI estimates if you had a complete pedigree going back all the way to the founding of the breed. So really, if you're only looking back five to 10 generations, that's better than nothing, but you're missing all of this inbreeding that is necessarily occurring because these dogs ultimately descend from a small set of founders. And this small set of founders is contributing real runs of homozygosity into the genes of these dogs. It just isn't being detected with the pedigrees. And finally, pedigree COIs are just estimates, right? So, the amount of genetic sharing you see between two individuals is not just the, “Oh, half of the genome came from this ancestor, half the genome came from this ancestor, half the genome came here, half the genome came here; let's multiply all those halves together, and that's the answer,” because we know that there's randomness in the process, so it actually depends on which chromosomes got transmitted to which ancestors at the DNA level for how much of the genome is actually identical by descent. So full siblings will always have the same pedigree estimated inbreeding coefficient, but the actual COI, when you estimate it from genetic data, is going to vary. Because again it depends on which chromosomes got inherited by each sibling. 


All right, so to conclude, when you're looking at a pedigree COI estimate, you're looking at a calculation of what the average COI would be in a dog that had all the known ancestral relationships recorded in the pedigree, assuming the pedigree was accurate and complete, and the founders of the pedigree were unrelated. The pedigree COI is the expected average value. Two littermates are always going to have the same expected value, but they're going to have different actual values when you look at the DNA, because of the randomness in the way the actual DNA was passed down from each ancestor. In contrast, genetic COI is based on analysis of runs of homozygosity, and then it's an estimate of the actual COI that the dog inherited, and we say it's an estimate because we're using the runs to infer identity by descent regions, but the DNA doesn't tell us exactly where identity by descent begins and identity by descent ends. Still, it has several advantages over the pedigree COI calculations, not the least of which is that with modern genotyping arrays, it's able to infer inbreeding from ancestors much further back than most recorded pedigrees. Because of this, genetic COIs tend to be higher than pedigree COIs for most dogs. They are also trying to measure the same thing: the fraction of the genome that's likely to be identical by descent in a dog. But the genetic COI just is able to look further back and is able to see differences between siblings when doing this. 


For the last portion of the talk, I want to talk about a genetic topic that's related to coefficient of inbreeding, and that's genetic relatedness. So, inbreeding, as we've discussed, is when an individual inherits two identical by descent chromosomal segments, and we measure it with the coefficient of inbreeding. Relatedness, on the other hand, is when two different individuals inherit the same identical by descent chromosomal segments. In this case, the individuals are related to each other because they have the same DNA they got inherited from a shared ancestor, but they aren't necessarily inbred. We measure relatedness by the coefficient of relatedness, or little r. For those who are geneticists used to thinking about haplotypes, we sometimes refer to this as haplotype sharing. So now, instead of looking at long runs of homozygosity, where the two segments of DNA in the same individual are identical, we're looking across individuals to find DNA stretches that are identical. If the stretches are long enough, we say that they are identical by descent, and they contribute to the coefficient of relatedness. 


So, in this case, Molly and Loki share an identical chromosomal segment in this region, and if we added up all of the shared identity by descent chromosomal segments that they share across their genome, we could calculate their overall coefficient of relatedness. Now, coefficient of relatedness can also be calculated from pedigrees. Identical twins are 100% related, parent offspring are 50% related, full sibs are 50% related, so on and so forth. Half siblings are only 25% related, and grandparents/grandchildren are only 25% related. Now the relationship for these fractions can actually be a bit different in different pairings, so we say parent offspring are 50% related, and they're basically exactly 50% related. So, if you look at all of these individuals in this pedigree, they all got exactly 50% of their DNA from one parent and 50% of their DNA from the other parent. We say full siblings are 50% related, but this is just an average expected value, because the DNA got shuffled with recombination, so they on average will have 50% of the same DNA, but it could be higher, could be lower. And same thing with, you know, grandparent/grandchild is expected to have 1/4 DNA similarity, but if you go through and you paint the genome, let's say two cousins, you can paint the fraction of the genome that came from a specific shared grandmother, let's say, and you can see that, “Oh, okay, it's random, it's about 1/4 but it maybe in one it's 20% and one it's 30%” and then if those cousins were to cross, you would get this: you could see where the fractions overlap and where you would potentially have inbreeding in the offspring, and that's how relatedness and coefficient of inbreeding are related, right? So, if you have this overlapping segment where two individuals got DNA from the same ancestor, then those two individuals have this potential to transmit it to the offspring, and there's a 50% chance that'll happen. So, in general, the coefficient of inbreeding of offspring is going to be half the coefficient of relatedness of the parents. Now I say, in general, because there are, of course, complications to it. So, if we actually look at actual coefficients of relatedness, when you have full sibling crosses in different populations (so here's a cattle data set, here's a sheep data set), we do see this distribution around 50% with some individuals having less than some individuals having more, but we also see this tail of some individuals having a lot more. More than can be explained by chance alone. And this is because the parents themselves might be inbred, so if the parents are inbred and related, then the coefficient of inbreeding in the offspring is going to be more than 0.5. So, in the extreme cases, if parents have a coefficient of relatedness of essentially 1, if they're essentially identical twins, obviously they're not gonna be identical twins, because they're different sexes and stuff, but if you created, if you had a selfing species, for instance, when you cross those two identical strains, you're going to have a 50% inbreeding coefficient. Whereas when you have inbred lines, so not only are the parents identical, but they're also inbred, then the inbred line produces an identical inbred individual with a coefficient inbreeding of 1. So in general, the coefficient of inbreeding of the offspring is half the coefficient of relatedness of the parents, but if the parents are inbred, it actually can be elevated a bit above that. 


So, why do we calculate relatedness in genomic data sets? Obviously, the first answer is: well, duh, we use it to figure out who's related to who, and that's definitely true. We do it in human data sets because we want to make sure we don't include really close relatives and certain genetic analyses. I was really excited to do—I had a shelter dog at home, so we started Embark. We started testing a whole bunch of dogs. This is my shelter dog, Penny. And after we had Embark for a couple years, a dog named Dharma popped up, that was her first cousin, which was super exciting. And doubly exciting, Dharma only lives a few miles down the road from us, so we had no idea that Dharma and Penny were related. They don't recognize each other. They weren't littermates. Dharma was born about a year after Penny, about six months after we adopted her, but it was still great to find a relative of Penny who happened to also share some of Penny's characteristics and behaviors. So, I don't know that dog breeders appreciated this, but most adopted dog owners, they adopt shelter dogs, have no idea what their dog's family medical history is. So finding relatives is not only fun, but it potentially gives some of us shelter adopter owners useful information about caring for our dogs.


Now, for breeders, we want to know relatedness, because we want to know if we're crossing highly related individuals that are going to yield litters with high expected coefficients of inbreeding, so instead of getting a list of close relatives, breeders get a pair predictor tool that uses the relatedness of two potential mates to compute the expected COI of the litters they would produce, so I'll go into this in a lot more detail in the second half of the webinar series, but the main thing to know about this expected COI calculation is that we prefer lower COIs because higher COIs are associated with inbreeding depression. 


Inbreeding depression can affect all facets of an organism's life, including growth, health, lifespan, and reproductive success. So, if we go back to the example of King Charles the Second of Spain, inbred humans are at risk for inbreeding depression. His parents had a pedigree-based coefficient of relatedness around 50%, giving him an estimated coefficient of inbreeding of about 0.25. If we look at pedigree databases in dogs, we can identify dogs that have high or low 10 generation coefficients of inbreeding, and we can see whether it's associated with longer or shorter lifespan. In Golden Retrievers, we see that every 1% increase in the COI in the 10 generation pedigree yields about a one month shorter expected lifespan. We also see inbreeding depression, if we look at the reproductive success in Golden Retrievers. Every 10% increase in inbreeding is associated with a one pup reduction in litter size. Finally, when we look at owner-reported Embark questionnaires, we see that the lifespan is shorter in big dogs than small dogs, obviously, but also that for any body size, the lowest quintile of inbreeding lived about two to three years longer on average than the highest quintile COI dogs. So, not only were these lower than average COI dogs living longer, when we also asked about their health, their owners were more likely to report them as being in excellent health. Interestingly, inbreeding levels vary a lot between dog breeds and also within dog breeds, so even dogs that are relatively outbred, like Beagles, have some individuals with very high inbreeding coefficients. In other breeds, all or nearly all of these dogs have COI values that are above the 0.25 that we estimated for King Charles the Second. 


I'll talk about this a lot more in the next webinar, but diversity and COI across breeds and within a breed is caused by many factors, in addition to inbreeding, like artificial selection for specific breed-defining traits, as well as loss of genetic diversity due to genetic drift in some breeds. If we look at the amount of genetic diversity across the genome in Labrador Retrievers, we see that the genome has relatively high genetic diversities, whereas in Doberman Pinscher's genetic diversity is mostly lower, and much of the genome is actually completely absent of genetic diversity, and so there's several large regions, like the beginning of chromosome 3 and the beginning of chromosome 31. Dalmatians are kind of intermediate, so they have mostly high diversity in a few areas with less diversity and a few areas with no diversity, and I'm circling three of those here because they lack genetic diversity due to selection. Chromosome 20 is the mid F white spotting locus, so that's obviously going to be fixed in Dalmatians. On chromosome 38 is the roaning locus, which is also fixed in Dalmatians. So Dalmatians carry the roaning gene, but their spot pattern is different because they also contain the spot modifier locus on chromosome 3, which is circle. So by breeding for this Dalmatian spot pattern, we've sort of fixed the genome in all three of these regions, and there's no genetic diversity until it got broken up by recombination further away from the selected locus. And the point is we don't want COI to be 0 in these breeds, because we want to keep these breed-defining traits. We don't want to outcross away from them, so instead we want to balance the various factors when making breeding decisions, including preferring crosses that would result in lower expected COIs over ones that would result in higher than expected COIs. There's a lot of variation in most breeds, and we can use that to our advantage to select for healthier litters. If we look at the distribution of inbreeding coefficients, let's say for Chows that we've tested at Embark, we see that as a breed they tend to have a bit lower than average inbreeding coefficient compared to most purebred dogs, and most commonly it looks like it's around 10–12% here in chows, but there's a very long tail to the distribution with some chows having inbreeding coefficients of 50% or more. By preferring crosses that result in litters with 10–20% litter COIs over those resulting in 40% or more, in balancing other factors when making breeding decisions, Chow breeders can make a real difference in the expected health of their litters. 


Okay, so that's all I was going to cover for today. I, of course, I'm going to have a second webinar, but I hope that there was enough there that people found it interesting, and if you have any questions, I want to stay on and give you time to ask them.


Nicole Engelman  41:48

Awesome. Thank you so much. This was such a great presentation. We do have a ton of questions. Please, as I ask them, let me know if it's just something we're going to cover in part two, and we can table it for now. So, we have maybe to start, a few questions about just practical applications of the COI in a breeding program. So, someone asked, when a breeder is evaluating a potential pairing, where should the COI fit in in their decision making relative to health testing, temperament, structure, and type?


Adam Boyko, PhD  42:21

For most conditions, you never want to cross carriers, because that's how you're going to produce affected individuals, and so that should be a primary decision. You just rule out a carrier/carrier crossing. Now, you can certainly breed carriers, you just want to breed them to clear individuals. So if crossing a carrier to a clear is going to give you a higher COI than crossing a carrier to a carrier, that doesn't matter. You don't want to cross carrier to carrier. But then beyond that, I would say, well, if you're looking at the difference between, you know, a 0.2 expected COI litter COI and a 0.4 expected litter COI, that's a pretty substantial difference, and so you would really want to see: is there another pairing that's nearly as preferable, you know, that is going to give you what you want, but at a much lower expected COI, or not? And if you can swap out even just some of them—you don't necessarily have to swap out all of them—I think overall you're going to protect the genetic diversity in your breed, and you're also going to protect the genetic health in your lines much better now. When we get down into, oh, this pairing would be a 0.21, and this pairing would be a 0.23, I would say that shouldn't be a really a price, like that's a tie breaker. Like, if you really can't make a decision between two dogs because their health testing is all the same, and you know the expectation about how they're going to perform is the same, you know, sure, use it as a tiebreaker, but that shouldn't be like a primary driving force.


Nicole Engelman  43:50

Thank you for answering that one. It sounds like we have a lot of breeders joining today that are coming from either really small breeds, rare breeds, breeds with limited gene pools. So, how should breeders that kind of fit into those categories think about the COI? Is it just the idea that keeping the COI low is genuinely difficult with these types of breeds?


Adam Boyko, PhD  44:14

Yes, so at some point the breed has lost so much genetic diversity that there's regions in the genome where there's no genetic diversity left, and so in every single dog, those are going to contribute to COI, and so if you keep a closed population, there's nothing you can do about that. You're essentially waiting for new mutations to crop up enough to differentiate it, and yes, we get new mutations every generation, but the odds of it happening in a particular spot are very, very low. So if there's nothing you can outcross to, then you're going to have high COIs, right? But by keeping track of litter COI, you can help slow down or stop the further loss erosion of COI. You know, in particular, what it does is, if an individual, if you've got a region of low genetic diversity, like there's only a few diverse haplotypes left in the population, it's almost, you know, ready to get wiped out—well, this expected COI is going to find those individuals and say, “Hey, if you did the pairing with this individual instead of that other individual that carries the common haplotype, you're actually going to get a lower COI, because that section of the genome is going to stay outcrossed instead of becoming incrossed.” So, it's a way to kind of slow down this inevitable genetic erosion that's going to happen in a small closed population.


Nicole Engelman  45:37

And just kind of going off of that, we had some people who are curious about your take on line breeding, and if there is a responsible way to do it, how would you balance the benefits with the risks?


Adam Boyko, PhD  45:48

I mean, there are probably more responsible and less responsible ways to do it. I think there's always a risk when you do it, and I'll get into this more in the next webinar. Homozygosity itself, inbreeding itself, isn't what causes the disease. It's the fact that the inbreeding occurred for a chromosomal segment that contained a recessive disease allele, that's the ultimate cause of the disease. Your risk is higher the higher your inbreeding coefficient is, but it's not a guarantee, and so I didn't go into it, but King Charles the Second of Spain had a sister who was fine and had kids, and so it's not a death sentence to have a COI of 0.25, obviously, but in normal dog data sets, within breeds and across breeds, we see, and in mixed breed dogs, so mixed breed dogs can be inbred as well, because you have backyard breeding, stuff like that. We see it in mixed breed dogs, just like we see in pure breed dogs. The more inbred ones, on average, have more problems, but that doesn't mean you can't have a particular dog with a high COI that is fine.


Nicole Engelman  46:55

I feel like this might have kind of been covered, but I did like this question as more of a general zoom out: What are the most common misconceptions you think that you see about the COI in the breeding community, just in general?


Adam Boyko, PhD  47:09

I mean, I get a lot of questions about why don't my Embark COIs match what the pedigree software is telling me. So, I'd spent a bit of time here talking about that, and hopefully that explanation made sense. There's, I think, a little bit of confusion about COI, you know, is a measure of genetic diversity, and it's not where high COI is low diversity, low COI is high diversity. It's a tool that helps retain genetic diversity, but it's not exactly the same thing, and so that's really what the next webinar is going to talk a bit more about, so really it's just “This is the expectation of the risk of inbreeding depression for an individual because this is the fraction of its genome that's exposed to the potential for inbreeding depression.”


Nicole Engelman  47:53

Great. The next set of questions we have are really about tools and resources, so if that's okay, we can get into that, or do you want to save that for the next?


Adam Boyko, PhD  48:02

We can get into it. Again, I'm not a dog breeder, so you know I'm familiar with tools I've developed and a few other tools, but I'm not like a master at doing this.


Nicole Engelman  48:09

No, definitely. I think it's helpful to have, honestly, an outside perspective. I see a lot of people just seeming to want to know where to start, and they were curious if there are tools or databases that you recommend breeders use to calculate and track COI over time.


Adam Boyko, PhD  48:26

Well, I mean, I made the pitch here, but of course I'm biased, because I started a dog DNA testing company that actually looking at the genetics and using that to calculate COI by looking at shared identity by descent regions is the best way to do it, because you get the clearest picture, and you can, importantly, differentiate between dogs that have identical pedigree COIs, but this dog actually is maintaining this rare haplotype that will contribute to diversity in the future, and this dog doesn't have it, because it just didn't get transmitted in that dog. And so just genetically testing, and then any potential crosses, so: Here are dogs that I'm considering for my breeding program, and now I know, in addition to what genetic defects they carry, what coat colors they carry, what traits I'm interested to carry, and all the health testing, everything else, but that I also know have an expectation of the COI of those litters that they would produce, so that I know I'm not accidentally line crossing or accidentally producing a, you know—it doesn't look like it should produce a high COI litter based on the pedigree that I have, but in fact they happen to inherit a lot of shared DNA from earlier ancestors, and this other dog that on the pedigree looks just as related, actually it's much less related, and so I can outcross. It's essentially outcrossing that it wouldn't have been able to do just based on pedigree estimates.


Nicole Engelman  49:40

And do breeders necessarily need, like, a deep source of pedigree data available, or do you think that anyone can?


Adam Boyko, PhD  49:47

No, I mean, that's what kind of democratizes it. I mean, if you, as long as you've got the genetic profile from the dogs that you're considering, you can calculate the COI, and so you're not reliant on pedigree software, and you get those other advantages that I talked about for using genetic COIs for consideration instead of just pedigree COIs.


Nicole Engelman  50:06

Yeah, I think that's a really nice takeaway for everyone here, because it does look like a lot of people are just looking for information on how to get started doing all of this. We did have one last question, because I want to be conscious of your time, and I know this one you might be a little bit biased about, but someone was wondering how we can make sure we're using the right genetic tests to get the most accurate numbers, and are all genetic tests created equal?


Adam Boyko, PhD  50:31

So, no, not all genetic tests are created equal. So, there are tests that aren't based on SNP arrays of hundreds of thousands of markers. So you know, let's say the previous generation used microsatellites, right, which are great. Like, microsatellites are great tools. We use them for DNA fingerprinting. You can use them for parentage analysis, although that's pretty much all shifted over to SNP arrays now, too, because it's just more efficient. But the drawback is you can only type so many microsatellites, so it's usually dozens, you know, maybe in big data sets, a hundred, but usually people are just typing dozens of microsatellites, and so you're not even looking at every single chromosome when you're calculating inbreeding. It's a genetic calculation, but it's going to miss most of the inbreeding in the genome, so you're getting this really rough kind of estimate, like, you know, chromosome 17 could be 50% inbred, but the odds of the one microsatellite and chromosomes 17 falling in and/or out, like it's either going to be zero or one inbred or outbred, right? And so you don't get that 0.5 thing, and you also have other estimates that you SNP arrays, things like heterozygosity, that yes, I mean they're correlated to COI, but it's not the same thing as COI, and so you'll get a lot more noise around your heterozygosity versus the actual COI estimates, so the genetic COIs are going to be more highly correlated to the pedigree COIs and vice versa than they are to genetic heterozygosity.


Nicole Engelman  51:54

Great, that was, I think, the last question we have time for. Thank you for answering all of those. I know there were a ton! 


Adam Boyko, PhD  52:00

Thank you, Nicole.


Nicole Engelman  52:00

And thank you for this amazing presentation. I just speak on behalf of our community. I think everyone loved it. I do feel like we've barely scratched the surface on this, which is a good thing, because I know we're having you come back on May 27 on Wednesday, so that will be part two of our discussion. Everyone, I know you all came prepared with questions. Please do so for the next one. I'm sure there'll be a ton of questions that everyone has.


Adam Boyko, PhD  52:26

I'll make sure to leave maybe a few more minutes next time too, in case.


Nicole Engelman  52:29

Yeah, maybe. I think everyone is super interested in this topic. So we're so grateful to have you cover it for us. And thank you, everyone, for joining us. And Dr. Boyko, we'll see you later this month.


Adam Boyko, PhD  52:39

Great, thanks, Nicole. Thanks, everyone.


Nicole Engelman  52:41

Bye. Thank you.

Share this article

Join our Good Breeder community

Are you a responsible breeder? We'd love to recognize you. Connect directly with informed buyers, get access to free benefits, and more.