Paper guide: Let's
focus this week on the biological results a little more than the
method. For the methodology, understand Figure 1 (in detail,
what all the symbols mean etc., you may have to do some
background reading on Hidden Markov Models). For the biology,
focus (as the paper does) on the relative amounts of conserved
DNA in different eukaryotic groups, and also think about how
this probably extends to bacteria, archaea, and viruses (not
covered in the paper). A second focus should be the fraction of
conserved non-coding DNA in different groups and what it means.
A hard question to consider: how should we define and measure
organismal complexity? This is a term used often in biology,
rarely with real meaning (often it means roughly "the closer it
is to human the more complex it is"). Things NOT to focus on: 1)
the data for nematodes (as the authors several times point out,
the two taxa available at that time are too divergent to give
reasonable results; unfortunately they include the results
anyway), 2) the section on secondary structure in non-coding
HCEs.
Paper guide: The primary paper uses the methods for dN/dS analysis developed in a series of papers by Yang and Nielsen (and often Goldman and a few other authors). In the Swanson paper, read the statistical analyses Methods section carefully and relate it to the Yang et al. paper (see below). In the main text and Table 1 you can discount the M3 vs. M0 tests (these are not very realistic models) and concentrate on the M7 vs. M8 tests, which are the most commonly used models. It is rather complicated, but try to understand all the numbers in Table 1, especially the Likelihood Ratio Tests (2 deltaL) and Parameter estimates. As for the biology, pay close attention to the Introduction and to the interpretations given in Discussion (you may have to read other papers or sources to understand fully why these processes might give rise to positive selection). Also think about other processes that might drive positive selection, and strengths and weakness of the dN/dS approach to detecting them.
The Yang et al. paper is long and very complex. You can ignore
all of the tests on real data (the Swanson paper is more
interesting) and most of the various model combinations the
authors discuss. Concentrate on the Theory section (up to
Statistical distributions, which you can skim), especially the
M7 and M8 models and the Discussion. We won't cover how specific
sites are identified (it is complex and the essence is simple -
it looks for sites where the maximum-likelihood dN/dS value is
high).
Swanson et al. (2001).
Positive Darwinian selection drives the evolution of
several female reproductive proteins in mammals. PNAS 98:
2509-2514.
Paper guide: There is
great interest in predicting which alleles segregating in
human populations are deleterious (i.e. cause "disease", or
more broadly cause phenotypes). Most such alleles are thought
to be neutral or nearly neutral (have little or no phenotypic
consequence). The Kircher paper takes a new and different
approach to this issue by using fixed nucleotide changes on
the human lineage as a sample of nearly neutral changes and
comparing them to all possible changes (simulated), which will
be a mixture of neutral and deleterious alleles. Focus your
thinking on the general issues - how this approach can inform
analysis of any genome population - and don't focus on human
disease per se. The paper spends most of its time on
additional efforts to combine approaches and datasets and
using machine-learning to evaluate the combined data to
predict human disease alleles. Try to understand this part of
the paper at a high level, but don't focus too much time
there. Class discussion will be focused more on general
population genetic issues - where mutations come from, how
they become fixed in populations, etc.
Unless you have a strong background in population genetics and
molecular evolution, I strongly recommend reading the Duret
overview of some of the key issues.
Here are a few questions to
think about: how can you explain the preponderance of diversity
shared by all groups (remember there were sometimes multiple
rounds of small founder groups)? Similarly where do you think the
population-private alleles come from? How can you explain the
groups that prominently have two colors in Figure 1 (e.g. Mozabite
and Hazara)?
Do take a pass at the Pritchard paper, in particular to
understand the K values and the methods used to cluster
populations, but the method is complex so don't get too bogged
down in the details. [Ignore the font changes here - my
HTML editor has a bug that is preventing me from changing it even
when I edit the HTML source code.]