Keywords: Design, sampling, data analysis, biology, medicine, health, agriculture, ecology, biodiversity, mathematics, statistics
Contents
3. Biometric Data Collection and Analysis
Summary
Biometry is a discipline devoted to the mathematical and
statistical aspects of biology. The benefits to mankind of biometrical
developmentsranging
from their applications in agriculture, and in animal and plant sciences, to
those in medical science, and public health
have
been enormous.
The interface between biology and mathematics presents many challenges. Fundamental, and important, research problems have surfaced here. This process continues, both because of the explosion of biological data with the continual development of new technologies, and because of the development of ever more powerful computers to organize and analyze the plethora of data. In biology, the challenges range from data at the molecular level, to the biosphere. In mathematics and statistics, the challenges range from application of existing methodologies to the development of new ones, tailored to the biological application, with the aim of giving both broader and deeper insights into biological data and biological systems.
Biometry is a large and complex field that arises from the application of statistics and mathematics to biology. All phases of research in biology, including design and data collection, analysis, and interpretation of results, depend on statistical principles and statistical methods. Discard any notion that biological statistics is all about hypothesis testing and p-values! The aim of a well-planned biological investigation is to gain insight into questions of scientific, biological interest. A well-developed biomathematical model that accurately describes the data aids in understanding what the data say, and so in making predictions and forming new questions. Unfortunately, many non-scientific, ad hoc, procedures are practiced under the banner “biometrics.”Two reasons have been given for the perpetuation of this state of affairs. Firstly, many biologists and medical researchers are trained without getting any real insight into the methods of science. Secondly, many editors trained in this manner will not accept papers for publication unless they follow these (often well-ingrained) ad hoc procedures.
Underpinning many areas of biometry is the mathematics of
probability. In particular, special stochastic models are often developed.
Examples include models in genetics, in particular in population genetics, in
epidemic theory and predatorprey
interactions. The type of model depends on the context. Sir David Cox has drawn
rough distinctions between purely empirical models and (at the other extreme)
“toy” models; in between lie the intermediate and quasi-realistic models. A
purely empirical model has no direct link with the underlying biological
process or corresponding interpretation of the parameters. An example of an
empirical model is the fitting of a curve to, for example, AIDS incidence data
observed over time. In a “toy” model, a highly idealized representation is used
to explore the particular circumstances under which a phenotype of interest
could be generated from simple starting assumptions. Examples include
biological models showing conditions for the extinction or explosion of
epidemics, or the extinction of species by competition. An intermediate model
is one in which some aspects of a complex biological process are represented,
with the objective of obtaining a formulation such that the resulting parameter
estimates do have a link with the underlying generating process. An example is
the back-calculation procedure used to predict HIV incidence from observed AIDS
incidence, assuming a particular formulation for the incubation distribution.
Such a model can produce a reasonable fit to the AIDS incidence data. A
quasi-realistic model involves complex processes in biological systems. Such
models are usually deterministic rather than having an explicit stochastic
component.
A natural question arises: when is the introduction of a stochastic element into a model likely to be crucial? For example, in epidemic models the deterministic model gives the corresponding stochastic means, but for small systems such a model may give a poor idea of the behavior of the sample paths. However, the often more biologically realistic, deterministic “toy” models (for which explicit solutions can often be more easily found) can be used with even more realistic and elaborate stochastic models in the interpretation of results from a complex simulation model. For example, the ratio of relevant response variables may be examined in comparison with those predicted by the “toy” model.
For data collection and analysis, the same fundamental principles apply to experiments, to observational studies, and to the secondary analysis of data collected (usually) for another purpose (such as for a disease registry). In essence, the aim of the study is to provide insight by means of numbers, and it is useful to distinguish three broad headings:
Collection of data.
Organization
of data.
Drawing conclusions from data.
In all types of study, the key initial questions are:
What units (individuals) should be
included?
What
properties should be measured, and how?
What interactions should be examined?
Essentially, in the planning stages, one needs to consider how to control the random error, and avoid systematic error (bias). There is an enormous literature on these issues. Basic principles of design need to be better understood. This is especially true in the laboratory sciences, where a widely-held but erroneous view is that refinement of laboratory technique is preferable to statistical methods for error control. A key element of the scientific method is replication. This needs to be better understood, and applied more in practice.
Good tabular and graphical procedures are invaluable, and have become easier to produce with developments in computer software. Certainly graphical procedures should be more widely used, both for exploratory work and in presenting the conclusions of more elaborate analyses. Also, a current research focus is the area of graphical models that allow interactions between parameters to be clearly shown, and complex structures can now be fitted relatively easily using modern computer power. In statistical analyses that are heading towards the “conclusions” stage of the study, the methods depend, at least in part, on an explicit probabilistic base. When choosing a model, the following should be considered:
The model should be consistent with
previous related studies of the topic.
If
possible, the model should establish a link with underlying substantive
knowledge.
The
model should be consistent with, or suggest, a process that might have generated the data.
The
parameters should have clear interpretations, and the error structure should be
such that measures of precision are meaningful.
The fit to the data should be
adequate.
These days there is an increasing need to relate the primary conclusions in different studies, including examination of the consistency of conclusions. In medical research this has lead to much interest in meta-analysis. Another fundamental concern for many scientists is the underlying “causal” process. Causality is a “slippery” notion, and a cautious usage is that strong evidence for causality can only come from the synthesis of different kinds of data. Closely related are the complex issues of generalizability and specificity, as well as the importance of any interaction/s that might be present, but has been assumed absent.
Concerning inference, there are many different approaches, ranging from pure likelihood to the Fisherian approach (with its emphasis not only on likelihood but also sufficiency, conditionality and ancillarity), and the Neyman-Pearson approach (with emphasis on power), to the Bayesian approach. The use of highly sophisticated models in the analysis is not always necessary. Both sensible statistics and sensible biology will prescribe the final form of the quantitative model. The underlying assumptions of any model invoked in the analysis must be carefully checked; if any violations occur, the biological significance of the information provided by these models will be greatly reduced. A pointer to the need to examine assumptions will occur when the data disagree markedly with expectations. Moreover, the best quantitative models are useless anyway if they have little relevance to the biological processes they were meant to describe. Certainly analyses should not be an end in themselves; rather they may provide a springboard to as-yet unanswered questions about the biological system being studied. Overall, the selection of an appropriate statistical model for the analysis will be iterative. Both biological and statistical principles are needed to define and refine the quantitative models used to describe the biological processes.
Risk assessment and management are important, and there is extensive discussion of these issues in respect of epidemiology, toxicology, and other topics. The role of judgmental probabilities in such situations is central. In general, individuals have little understanding of extreme probabilities, and their evaluation. Quality control and process improvement methods are also used in various biometric applications. For example, they can be used in multicenter clinical trials to provide quality medical evaluation for the final evaluation.
The problems and questions faced by real-world applied biometry are widespread and far-reaching. How do some birds learn to navigate so well? What factors influence the length of time individuals spend in institutions, like hospitals or nursing homes? Can a particular fish species survive being caught and returned to the ocean? Which histological changes predict cancer at a certain site? No essay, like this one, or set of theme headings, can properly capture the richness, breadth, and depth, of biometrical data and its analysis.
There is no shortage of interesting new ideas and challenging problems in biometry, with many of these stemming from the relatively large datasets that are now proliferating. Collaborations between biologists and biometricians are essential in developing biometrical modeling methods for research in biology. In particular, many current and future challenges are being motivated by questions in molecular biology, genomics, proteomics, and molecular evolution. These will require the development of new techniques and theories.
Much of the following is based on material that is given in the bibliography. This overview is not a comprehensive overview of biometry, but offers a broad-brush picture of the past, present, and future of this fundamental discipline that underpins so much Life Support Systems knowledge and ongoing research.
Since the seventeenth century, biological phenomena, like
mortality and morbidity, have been the central concern of those who collected
and analyzed statistical data. John Graunt (16291674)
and William Petty (1623
1687)
were two pioneers of this time. During this period, the mathematical theory of
probability developed from interests in games of chance, and gambling, and
major pioneers were Pierre de Fermat (1601
1665),
Blaise Pascal (1623
1662),
and Jacques Bernoulli (1654
1705).
Abraham de Moivre (1667
1754)
also was a pioneer in probability theory. He discovered the approximation of
the binomial distribution by the normal distribution, as well as investigating
mortality statistics. In the eighteenth and nineteenth centuries, the stimulus
was astronomy: leading pioneers included Pierre Laplace (1749
1827),
and Karl Gauss (1777
1855),
who realized the importance of random errors in observations, and developed the
method of least squares (that underpins regression). Another astronomer,
Adolphe Quetelet (1796
1874),
applied statistical methods to problems in biology and medicine. Pioneers in
epidemiology also emerged. Louis René Villermé (1782
1863)
correlated the variation in mortality he observed in data collected in
1883)
studied the distribution and determinants of health disorders in English
populations: his studies on mortality differences between different occupations
helped in understanding industrial hazards. A major epidemiological discovery
of this period was John Snow’s 1854 demonstration, using numerical arguments,
that cholera was a water-borne disease.
Francis Galton (18221911),
a cousin of Charles Darwin, made a substantial input to the birth of biometry.
Galton found
1910)
pioneered the compilation of relevant vital and medical statistics, accompanied
by vivid and revealing graphic representations.
The mathematical work of Karl Pearson (18571936),
and his colleagues like Raphael Weldon (1860
1906),
laid the foundations for modern biometry, and influenced many, including the
pioneer medical statistician Major Greenwood (1880
1949).
The dominant figure of twentieth-century biometry was Ronald A. Fisher (1890
1962),
whose vast contributions included the development of analysis of variance,
maximum likelihood methods, and experimental design. Problems in eugenics and
in plant breeding motivated Fisher’s statistical work. Work after the Second
World War saw a rise in epidemiologic studies focusing on associations between
a wide variety of factors and disease, like smoking and lung cancer. In the
Many diverse problems in evolution and genetics have had a fundamental
influence on both probability theory and statistics. Galton and Watson (1874)
founded the theory of branching processes as a consequence of their
investigations of the extinction of human family names. In the mid 1920s,
McKendrick and Kermack developed non-linear birth and death processes in
answering epidemic theory problems. The work by William Feller (19061970)
on stochastic processes was partly motivated by population genetics problems.
Counting process models have been developed for studying patterns of arrivals, and interactions of nerve impulses from different neurons. Markov processes have been used in analyzing membrane channel data, studying the kinetic behavior of ionic channels, and understanding DNA damage caused by ionizing radiation. Stochastic differential equation models have been used for investigating the depolarization of the membrane potential of spatially distributed neurons. The stochastic nature of the measurements has resulted in new developments in stochastic integration and differentiation, and growth of this mathematical field has been stimulated by neurobiology.
In summary, mathematical and statistical techniques have grown in importance over the past century, as has the way in which these methods have been used in biological research and practice. The “green revolution” in agriculture would have been impossible without these tools. Modern medicine and public health practice depend upon carefully designed and interpreted clinical trials, and upon massive observational datasets.
Finally, the rapid increase in computer power in the modern era has seen development and implementation of new ideas that have made a huge impact on biometrical methodology, both for the design of data collection and for analysis. Insightful graphical procedures have become easier to implement too, and good analyses today are accompanied by relevant graphs.
3. Biometric Data Collection and Analysis
![]()
3.1. Experimental Design
Experimental design, particularly its historical application to agriculture, has been an important tool in the advancement of biometry. Sir R. A. Fisher, who established the Statistical Laboratory at Rothamsted Experimental Station in the early 1920s, published two articles on crop variation that led to a worldwide revolution in the technique of agricultural trials. This is widely acknowledged as the starting point of research on experimental design.
The basic concepts to increase the accuracy of an experiment were formulated over the next two decades, namely:
1. To increase the size of the experiment.
2. To refine the experimental techniques as much as possible.
3. To select and organize the experimental material to minimize experimental variability.
The key elements of this last concept include the blocking of experimental units into groups that are as homogeneous as possible, and the use of covariates. Technically, improved experimental design involves development of increasingly sophisticated experimental layouts, along with corresponding methods of analysis. Good experimental design requires an understanding of the objectives, and the nature, of the experimental units. Randomization in the assignment of treatments to experimental units is fundamental to reducing possible biases from other sources of variation, that are either unrecognized or of no importance to the question/s under investigation. Increased computational power in recent times has seen the development of more complex designs, along with computational algorithms. As well as in crop research, experimental designs are used in animal research (for example, into dairy animals and pigs), and in evaluating drugs and other medical treatments. Classical designs for experiments include factorial experiments, fractional factorial designs, balanced incomplete block designs, Latin square designs, and Lattice designs.
For the basic analysis of such designs, Fisher introduced what is termed the analysis of variance (ANOVA). This provides a worktable for evaluation of the null hypothesis of all levels of categories of treatment having the same effect on the (continuous) outcome measured. Under the assumption of normality, relevant ratios of the mean squares (sums of squares divided by their degrees of freedom) can be shown to follow well-known distributions, against which values from the relevant test of the null hypothesis can be evaluated. There are strong links between these models and linear regression. The approach is readily extendable to accommodated concomitant variables (ANCOVA: analysis of covariance). These days the data can be clustered, and random effects models, as opposed to fixed effects models, can be utilized. Robust approaches to guard against inappropriate assumptions have also become increasingly popular.
Models of experimental data are often for prediction purposes: for personal ozone exposure assessment, for example, or for estimating seedling responses to environmental effects. When making inferences over some population of conditions, it is important that the increased uncertainty that naturally accompanies the broader inference is reflected in the associated measures of confidence. Assuming, say, that a fixed effects model generally overstates the level of confidence in the estimates, it is important that the experiments have randomly sampled the reference populations of conditions. This can be a problem if, for example, research sites are chosen for specific reasons, such as convenience. Environments sampled can almost never be regarded as truly random, and one hopes that the environments encountered over the experiment’s timespan are reasonably representative. To the extent that the sample of environments is not representative, there will be bias in the estimates.
Good design of experiments, and corresponding good data analysis, remain fundamental to good research. The number of important areas that benefit from good design are always growing. Today they include topics from diverse research areas: evaluations of chemical pollution; of effects in social experiments, such as offering economic incentives to see if there are effects on lengths of stay in nursing homes; of the feeding behavior of birds, to gain insight into their spatial association learning; and research on the ozone level. Often, cost considerations can be accommodated. For example, recent research has shown diagnostic tests for disease prevalence on pools of serum samples can, when properly designed, reduce cost and yet increase precision. As computer power has been increasing, so we can contemplate greater complexities in our modeling and data collection, and can better combine fragmentary data while simultaneously analyzing multivariate responses.
3.2. Sample Surveys
Sample surveys are widely used by biometricians, particularly in the medical field: for answering questions about community health issues, for example. The most compelling argument for using a sample survey, rather than a census, is that a properly conducted survey can provide more accurate results faster, for the same cost, and with an increased scope. Such a sample can increase accuracy by reducing the bias and increasing the precision of the results, and effort can be put into minimizing the non-response. Also, more questions can be asked, or more detailed measurements taken. Further, it is possible to study particular subgroups, such as older persons or ethnic minorities, in greater depth. Many epidemiological studies use sample surveys, and clever design can increase power and reduce bias in studies of risk factors for disease onset and progression. This is mainly achieved by oversampling groups at high risk.
The incorporation of probabilistic sampling into the selection of
units in sample surveys came in the 1930s, after the realization of the
importance of randomization in experimental design. The catalyst was the
increasing need for reliable estimates for use in policy making. In the
Today, data from both national surveys and community population studies are being used to understand causes and prevention of disease. By way of example, most of what is known about the risk factors for coronary heart disease comes from the results of long-standing community studies, such as the Framingham Heart Study and the MONICA (monitoring of trends and determinants in cardiovascular disease) project. Large-scale health surveys are done in the areas of health, nutrition, hospital discharge, ambulatory medical care, nursing homes, medical expenditure, long term care, longitudinal mortality study, and world fertility. There are also surveys from time to time on other health issues, such as heart health, drinking behavior, and mental health, to name a few, and on public attitudes to health issues (for example, towards workplace smoking restrictions).
A clear statement of the objectives of a sample survey, along with the design specification, is a fundamentally important part of the planning of a study. Surveys are usually of one of two types: descriptive, and analytic. Most large-scale surveys are of the descriptive type. Fundamental to the survey design is the notion of a random sample: in particular, for simple random sampling in a finite population of size N, the aim is to select n units in such a way that each possible combination has an equal chance of being chosen. Survey statisticians use more complex designs that are based on ideas concerned with, but not limited to, stratification and clustering. Use of stratification involves division of the population into non-overlapping subpopulations, and the sample is drawn independently from each stratum. There are many reasons for taking this approach: an important one is the ability to make estimates with a certain level of precision for subgroups of the studied population. Sampling in which units sampled are chosen in groups, or clusters of smaller units, is used for two reasons. First, a sampling frame (list of population units to be sampled) may not exist. A sampling plan could be designed, for example, to estimate the number of trees in a given geographic area. However, it may be possible from, say, maps to divide the region into subregions, and these subregions then being selected for study. The second reason is to cut costs. If conducting a survey of general practitioners, for example, rather than traveling the country to interview a random sample, it will be more practical to select a relatively small sample of geographic areas, and conduct interviews with physicians selected from these areas. The choice of cluster size involves balancing costs against precision. Also, analysis of cluster sampling data requires careful consideration of the intra-cluster correlation.
Besides errors arising from the sampling process, there are four common, major non-sampling errors. First, the sample may not adequately cover the universe of units of interest. Second, it may not be possible to measure some of the units in the population chosen for the sample: some individuals may not be able to be located, or may refuse to answer some or all of the questions. Third, the measuring device may not be able to determine accurately the questions being answered (food consumption questionnaires, for example, suffer from this problem), or some individuals may not understand the question. Finally, errors may arise in the recording, coding, editing, and tabulation of the data. There is a considerable volume of literature addressing these errors.
Finally, analysis is an important component in sample surveys, and complex procedures have been developed to fully extract information from what today have generally become very complex designs. Consideration of such matters is beyond the scope of this general-level article.
3.3. Graphical Displays
“A picture is worth a thousand words” is a well-known saying, and in our context this could be expanded to include “and a hundred biometrical analyses”! Statistical graphics are a critical component of modern data analysis, playing an important role in every stage. They are used in exploratory data analysis to determine the broad features of data and relationships between variables, as well as for diagnosing model inadequacies and determining model refinements. Further, they are used for data summary, data storage and retrieval, and for compact, forceful reporting of results. The best graphics transfer complex information simply, efficiently, unambiguously, and without distortion.
The necessity for statistical graphics in data analysis is
provided by Anscombe’s famous regression data (The American Statistician,
27, 1721)
shown in Figure 1, depicting four datasets for which the usual standard output
from a typical regression program is identical (including producing the same
line y=3+0.5x). Obviously the actual relationships between the
variables are very different, and in only the first case is the fitted line
sensible. This example shows how graphical tools can facilitate critical
thinking about data.
Figure 1. Four Scatterplots of Regression Data
3.4. Multivariate and Multidimensional Analysis
Biometrics, along with psychometrics, has made notable contributions to this area. In multidimensional analysis, samples are represented in a multidimensional space. A primary interest is to find visualizations of the multidimensional space, with an emphasis on two- (or three-) dimensional representations that approximate the “true” multidimensional space. There are strong links here with modeling approaches to data analysis. Some of the topics that fall under this heading are multidimensional scaling (including principal components analysis and biplots), models for two-way tables (including correspondence analysis), discrimination, and classification.
3.5. Linear Models and Generalized Linear Models
The term “linear model” refers to a model formulation that relates a response variable Y to a set of explanatory variables X1, X2, … Xk, where the basic relation is written
y = β0 + β1x1+β2x2 + … +βkxk + e
where y is the observed value of Y, corresponding to the observed set of explanatory variables x1, x2, …, xk, and the βs are regression coefficients to be estimated. It is assumed that E(e) = 0. The linear part of the model refers to the βs; the x variables may be discrete or (transformed) continuous variables, polynomials, or interactions. It is usual to write this equation in matrix form. A further common assumption is that all e values have the same variance, σ2, and that the covariance of all pairs of e values is zero. Estimation of β does not require any further distributional assumptions. To continue with hypothesis testing and interval estimation, however, it is customary to assume that the e’s are normally distributed.
Many standard statistical methods fall into this class of models, including regression (both simple and multiple linear), analysis of variance (ANOVA), and analysis of covariance (ANCOVA). It is interesting that R. A. Fisher did not base his original formulation of ANOVA on this model, but rather on the direct partitioning of sums of squares.
This model is a special case of a generalized linear model that refers to a regression model, relating a function of the conditional mean of a response variable, Y, to a linear function of the explanatory variables, x. Writing E(Y|x) = μ, a link function g is invoked, and the model assumption takes the form
g(μ) = β0 + β1x1 +β2x2 + … + βkxk.
The error distribution is also generalized, usually in such a way as to complement the choice of a link function. This leads to a very broad class of “regression” models, including the following two widely-used special cases.
Supposing Y is a Poisson random variable, and the link function g is the log function, then this can be shown to give what is termed the log-linear model. This is widely used, especially for the analysis of categorical data.
Suppose that Y is a binomial random variable (that is, the response variable is dichotomous) with π(x)=E(Y|x), and the link function g is the logistic function, log(π(x) / (1−π(x)), then this can be shown to produce what is known as logistic regression. Other link functions are often used for binary data, including the probit link, where g is the inverse of the cumulative distribution for a standard normal distribution. In applications involving discrete time survival data, the complementary log-log link (log(-log(μ)) has connections to proportional hazards models, and so is frequently used.
In many applications there are large measurement errors in some of the variables. An example is in ecology, where variables of interest are usually impossible to measure exactly due to time and cost constraints. In such situations, measurement error/errors-in-variables techniques have been developed.
Assumptions should always be evaluated, and the approach modified accordingly if needed. A standard assumption is one of additivity of the explanatory variables, when actually there is interaction between these variables, and, moreover, this interaction may vary dependent on other factors. An example might be response to combination drug therapy in disease treatment in the presence of other confounding variables (for example environmental factors, like diet).
Computer power has changed the complexity of models that can now be routinely invoked in standard analyses, compared with what was possible just a few decades ago.
3.6. Categorical Data Analysis
Categorical variables separate observations into groups, within which members share a common trait. This may be a nominal attribute, or a level of an ordinal scale, or a numerical value derived from an interval or ratio scale. Finer classifications can be obtained using combinations of several variables. Categorical data analysis refers to the response variable/s being categorical, and draws inferences from probability distributions of random category counts, or functions of them. A contingency table is a multi-way cross-tabulation of non-overlapping counts, where each dimension is categorical in nature.
Probability modeling is predominantly based on the Poisson distribution, and can be regarded as a special case of generalized linear modeling. However, other approaches are widely used, including weighted least square functional modeling, generalized estimating equations, and generalized linear mixed models (when there is extra-Poisson variability due to sample heterogeneity), as well as conditional logistic regression (when the response is two categories), and Bayesian inference. If numbers are small, and asymptotic results are not appropriate, then exact procedures have been developed, the best-known being Fisher’s exact test.
3.7. Survival Analysis and Risk
Survival analysis is the study of the distribution of life times (that is, times from an initiating event such as birth, start of treatment, or entering an institution, like a hospital or nursing home to a terminal event, such as death or relapse, or departure from the institution). A major feature of survival data, particularly in its medical applications, is the presence of incomplete observations; often it is only known that the events of interest occurred after certain points in time, and this is referred to as right censoring. Survival analysis as a mainstream biometric discipline is relatively recent, although there is a long history (going back to the seventeenth century) of the use of life tables in demography and actuarial studies. Survival analysis is use in non-medical areas too, such as size regulation: for example, determining survival when undersized fish are caught and then returned to the ocean. In medical applications, where all survival and censoring times are known precisely, a major breakthrough (1958) was the Kaplan-Meier estimator. This eliminated the need for approximations to grouped times, by shrinking the observation intervals to include at most one observation per interval.
In 1972, Cox revolutionized survival analysis by the introduction
of the “semiparametric” regression model for the hazard (now called Cox
regression), depending in a nonparametric way on time and parametrically on the
covariates. Although this approach now dominates survival analysis, other
regression models, such as parametric regression models, also play important
roles in practice. These include accelerated failure-time models, and
Models for survival data may be considered a special case of a multistate model: namely a model with a transient state “alive” (0), and an absorbing state “dead” (1), and where the hazard rate is the force of transition from state 0 to state 1. Such models may be conveniently studied in the framework of counting processes. One important extension of the two-state model for survival data is the competing risks model, with the transient alive state and a number, C, of absorbing states corresponding to death from cause c, c=1, …, C. Another important multistate model is the illness-death model with two transient states, say healthy and diseased, and one absorbing state, dead.
Risk is the probability that an individual without the
characteristic of interest to the study, such as disease, will develop this
characteristic over a defined age or time interval. Let Xi
represent the survival random variable for the ith individual. Then
mathematically the survival function can be represented by Si(t)=Pr(Xi>t),
and the hazard function by
λι(u)=limΔu>0Pr(Xi≤u+Δu|Xi>u)/Δu and so Si(t)=exp(-Λi(t)), where
Λi(t)=∫λι(u)du is the integrated hazard over [0,t).
So the risk for the interval [0.t) is calculated from 1-Si(t);
for small hazard rates, this expression is approximately equal to the cumulative
hazard.
3.8. Meta-Analysis
Meta-analysis is the systematic and quantitative review of the results of a set of individual studies for the purpose of integrating their findings, as well as explaining their diversity. The statistical basis of meta-analysis is now well developed. It has had a major impact on medical science over the past decade or more, and has formed the basis of evidence-based medical practice. One major problem is publication bias, which arises whenever the probability that a study is published depends on the statistical significance of its results, and methods for detecting/adjusting for this bias continue to be developed. More generally, relevant approaches to modeling and analysis form an active research area, including studies where the units are clusters rather than distinct individuals. The most challenging issues, however, concern which material to include in the synthesis.
Other areas of application are emerging, including meta-analysis of studies addressing genetic associations for various outcomes of disease. It has been shown that significant between-study heterogeneity is frequent, and that the results of an initial published study correlate only moderately with subsequent research on the same association. This discrepancy may be due to either bias and/or real population heterogeneity.
3.9. Bayes and Empirical Bayes
The Bayesian method, or paradigm, is a different approach to all statistical inference. In the Bayesian approach one wishes to calculate the probability of θ, the parameter of the distribution of the random variable X, given X=x, and by application of the Bayes theorem we know
p(θ|x) = p(x|θ) p(θ) / ∫ p(x|θ)p(θ)dθ .
The distribution p(θ) is called the prior (that is, the distribution of θ before any data are available), and p(θ|x) is termed the posterior distribution. In this approach one is therefore querying the probability that the hypothesis is true given the data, compared with the classical approach that asks what is the probability of the observed data assuming the hypothesis is true. Much debate over the past several decades has centered on what assumptions are reasonable concerning the prior distribution, so that one can calculate the probability the hypothesis is true given the data.
Implementing this approach generally requires extensive
calculations, which now are increasingly feasible as readily available
computing power has increased. Today biometricians are finding that Bayesian
methods, such as Gibbs sampling, and Markov chain
The controversy concerning the dependence of Bayesian analysis on a prior distribution for the model parameters, where the parameters for the prior (or some later-stage prior in a hierarchical model) are assumed known, can be assuaged by using the observed data to estimate these (final stage) parameters, or to estimate the Bayes decision rule directly, and then proceed as though the prior were known. This is known as empirical Bayes (EB). Both Bayes and EB methods are increasing in popularity amongst biometricians. Today Bayesian methods are being used in many areas, including spatial statistics, sequential analysis, model choice, experimental design, and sample-size estimation for clinical trials.
3.10. Computer-Intensive Biometrical Methods
The term “computer-intensive” was first applied to statistical
methods in connection with bootstrap techniques. It is now used to describe
techniques that depend, in some essential way, on the availability of
high-speed computation. Other examples include: nonparametric regression
(smoothing techniques); classification and regression trees; Markov chain
3.11. Nonparametric Methods
Statistical procedures developed early in the twentieth century tended to rely on underlying distributional assumptions, particularly that of normality. How justified these procedures were outside the distributional assumptions was found to depend on context. The first procedures that were found to be valid when these distributional assumptions did not hold included the sign test (whose origins can be dated back to 1710), and Spearman’s correlation. This was followed by development of procedures based on the ranks of the observations, and on what are referred to as counting statistics. The descriptor “nonparametric” is misleading, and these methods should be more clearly referred to as “distribution-free.” One of the areas where these methods have played a major role in the analysis of medical data has been in the analysis of censored data, particularly in survival analysis. Also, nonparametric methods of regression have been widely studied, particularly over the past decade. The popular smoothing estimates are constructed using either kernel regression or spline smoothing. Where there are several explanatory variables, the model is referred to as a generalized additive model, and can be considered as a nonparametric version of generalized linear models.
3.12. Time Series
A time series consists of values of a variable, usually recorded at regular intervals over a long period of time. Such data arise a lot in biometric studies. Some examples are hourly blood pressure readings, daily rainfall, weekly admissions to hospital, monthly mortality rates for a particular disease, yearly counts of a certain animal species, and animal body temperatures. Such data usually require special analysis methods because of serial correlation, namely neighboring measurements in a time series being correlated (usually positively) with each other, with this correlation declining the further apart the observations are in time.
Careful consideration of study design is a critical initial step in the model-building process. Essential design features include definition of the biologically meaningful variables, identification of the experimental unit, choice of appropriate sample size, selection of appropriate sampling interval, and choice of sampling period length. These all require fundamental knowledge of the biology of the system. Ultimately, the characterization of the measurement series depends on both the hypotheses to be tested and the assumptions of the underlying mechanistic model.
Following data collection, an essential first step is to plot the observations. Many time series can be regarded as a mixture of the following components: a trend; fluctuations about this trend; a deterministic cycle (such as a seasonal component); a residual or random effect. A common assumption is that of stationarity, namely that the probability structure of the series does not change with time.
Modern methods for the analysis of time series can be considered as being either in the frequency domain or in the time domain. Frequency-domain methods are used primarily to identify oscillations that explain a large proportion of the variance in the series, and Fourier analysis ideas and spectral analysis are widely used. Time-domain methods are based on direct modeling of the lagged relationships between each point and its past. An important first step is the correlogram, a plot of the lagged correlations of a series against lag size. The models most widely used are ARMA (autoregressive moving-average) and ARIMA (autoregressive integrated moving-average) models, whose parameters are usually estimated using likelihood methods.
Structural time series models are established in terms of components, like trends, seasonal effects, and cycles that have a direct interpretation. The coefficients in the trends terms are assumed to evolve over time as stochastic processes. The current, or filtered, estimate of the trend is obtained by first putting the model in state space form, and then applying the Kalman filter. The models can be nonstationary, and the seasonality and cycles can also be treated in a stochastic way. When the underlying data are non-Gaussian (for example, Poisson for count data, or binomial for qualitative data), the Kalman filter usually is not appropriate, and several alternative filtering and estimation approaches have been developed.
3.13. Longitudinal Studies
Longitudinal data is repeated measurement over time of individuals, or other units of observation. These repeated measurements may be obtained in a few seconds, or over many years of observation. Such data create special challenges, and special opportunities, for the biometrician. Longitudinal studies create opportunities to study individual patterns of change over time and conditions. These patterns yield estimates of the rate of change with time, age, or disease condition, for example, that are free of confounding effects due to factors that vary between individuals. Longitudinal data also provide more precise estimates of the rate of temporal change, or the effects of covariates on that rate of change, than would be provided by an equal number of observations on different individuals. Longitudinal studies have been an important part of documented health research for more than a century, based on studies of normal growth and development. Recently, more intensive interest in development, ageing, and changes over time in risk factors for chronic disease, has stimulated more intensive research in both design and analysis for longitudinal studies.
Early developments in methods for analysis of repeated measurement data were based on developments from linear models, and involved random effects and growth curve models, and mixed models for unbalanced and incomplete longitudinal data. In the 1980s, generalized estimating equations for such data and related work provided an integrated approach for data that could be either continuous or categorical (including binary and count data). Extensions to time-dependent covariates have been developed. Unbalanced and incomplete datasets can still play an important role in the interpretation of analyses, particularly if the missing data mechanism is non-ignorable (in other words, not missing at random).
Longitudinal data can arise from either surveys or experimental designs. When the research involves investigation of the changes of the response of individual subjects (who may, or may not, be grouped) to changing conditions (repeated measurements), principles of study design are well established. On the other hand, when the research involves modeling individual patterns of change over time, usually in different groups (growth curve studies), there are relatively few studies on optimal design. There is also literature reviewing the relative advantages of cross-sectional and longitudinal studies. Finally, the importance of rigorous measurement protocols, to reduce measurement error in longitudinal studies, is noted.
3.14. Spatial Analysis
Spatial statistics refers to the collection of statistical methods in which spatial locations play an explicit role in the analysis of data. Recent advances in computing have helped the development of more realistic models for the analysis of such data. It is useful to distinguish between continuous spatial variation, discrete spatial variation, and spatial point patterns. Modeling of continuous spatial variation refers to the analysis of a phenomenon in which the random variable Y(x) is in principle obtainable at any location x within a (usually two-dimensional) spatial region R. The variable may be continuous or discrete, such as concentrations of a pollutant at location x in a field, the number of organisms of a particular species found in a core sample centered on location x, or the presence or absence of vegetation cover at x. Discrete spatial variation arises for a random variable Yi that is associated with each of a finite, or countably infinite, set of fixed spatial locations xi. For example, Yi might correspond to the height of an individual tree in a plantation, or presence/absence of disease in that tree. For spatial point patterns, the locations xi themselves constitute the data, and are assumed to be generated by some underlying stochastic mechanism. Examples include the location of trees in a forest, or of cell nuclei in a microscopic tissue section.
Spatial statistical methodology has an extremely wide range of applications in biometry. Besides analysis of agricultural field trials, and of remotely-sensed images, models have been developed in geographic epidemiology to produce disease atlases; in environmental epidemiology, and in ecology, as well as soil science; and in fisheries (for example, estimating scallop abundance), to name a few other areas. Areas like neuroanatomy have provided motivation for development of three-dimensional models. Although replication is fundamentally important, unfortunately it has received comparatively little attention, especially in ecological and epidemiological applications, making it impossible for many studies to give firm reliable and recommendations (despite their claims to the contrary).
3.15. Image Analysis
Images to be analyzed in biometrics may come from microscopy, medical scanning systems (such as PET, positron emission tomography, and MRI, magnetic resonance imaging), electrophoresis, microarrays, and photography (used, for example, in agriculture and plant science). Although research on computer-based image analysis commenced in the 1960s, progress has not been as rapid as originally envisaged, as encapsulating what we “see” is not so easy! The ultimate aim of image analysis is usually to extract quantitative information which may, for example, be in the form of binary presence/absence categories, or measures of object location, length or area, or shape statistics. Image analysis methods are an eclectic collection of techniques, including linear and nonlinear filters, mathematical morphology, syntactic pattern recognition and computer vision, and Bayesian image analysis. Improvement in the power of image-analysis techniques has tracked improvements in computer power. Currently much empiricism prevails, but in time there should be convergence to a more systematic approach. Further, development of new technologies will offer new challenges to develop appropriate methodologies.
4.1. Agriculture
The use of biometrical techniques in agriculture research goes back many years. Moreover, agriculture was one of the areas for which challenges in data analysis spurred development of many modern analysis techniques. In particular, the technique of ANOVA was developed by R. A. Fisher at Rothamsted to analyze results from agriculture experiments. The earliest paper describing what may be considered a statistically-designed experiment can be traced back to Cretté de Palluel in 1788. It was concerned with an experiment on fattening of sheep in which, from four different breeds, a set of four animals were chosen and fed on four diets, one of each breed per diet, and the animals were killed at four-monthly intervals. The experiment can be regarded as a 1/4 replicate of a 43 factorial, or a 4x4 Latin square.
From the 1920s onward, agricultural journals have a long history of statistical writing. Further, courses on statistical methods applied to agriculture methods have been taught since 1915 (Iowa State College, USA). Since the 1930s, agricultural biometric methods have been greatly extended, by the introduction of new techniques, and today their use is worldwide. These techniques were initially employed in agronomy and crop husbandry, but similar principles were soon applied to experiments with animals, despite their usually greater expense and difficulty.
The most commonly used experimental design for agricultural field trials is the randomized block design, in which the area of land available for experimentation is divided into blocks of (hopefully) uniform soil conditions, and the blocks are subdivided into plots to which treatments are applied. (In fact, the terms “block” and “plot” in the experimental design literature reflect their source in agricultural studies.) The classical standard technique for analysis is the analysis of variance, to assess whether the variation among a group of treatments is greater than would occur if all observed effects were due to chance. Today much more complex models are often used.
Having done the experiment, the first task is to decide which variables are of interest for analysis, and to look for anomalous values that may sometimes lead to discovery of quite unexpected values: these may sometimes lead to discovery of quite unexpected effects. In some situations formal analysis is not necessary. Today, there is scarcely an agricultural experimental station in the world that does not use biometric techniques. Further, it has been found that methods developed in temperate climates have been useful in tropical developing countries.
Besides experiments on many types of crops, and in the areas of animal husbandry and disease and pest control, surveys using statistical methods have been conducted on a wide range of temperate and tropical practices in agriculture and animal husbandry.
Field trial techniques are fundamental to crop research, and although randomized designs (discussed above) are widely used, there is a role for various types of systematic designs that are analyzed using model-fitting techniques. Soil heterogeneity has long been recognized as a problem in field trials, and recent methods have been based on nearest neighbor analyses and Bayesian techniques. The responses of crops to fertilizers have long been studied, and response surface experiments and analyses have been developed. Plant breeding and variety trials can pose many biometric challenges. In particular, there can be problems with the combination of the results when variety trials have been repeated over several seasons and sites, and incomplete datasets can be problematic. Besides estimating yield, many crop trials are concerned with the amount of clean crop obtained by controlling pests and diseases. Study of the pest or disease organism may be of importance in its own right.
Since the late 1970s there has been formal research into intercropping, namely the growth of two or more crops together on the same area of ground for part of their life cycles. Intercropping systems are common in the tropics and semi-arid tropics, where conditions are harsh, farms are small, and it is difficult to produce sustainable yields. Initially statistical techniques were borrowed from studies on interplant competition. Demand to meet the specific requirements of this type of data quickly saw the development of novel methodology, however. For example, the Land Equivalent Ratio index allowed comparisons of intercrops with sole crops for systems with different proportions of land devoted to intercrops.
Besides animal research, biometric techniques have been developed for aquaculture and fisheries studies. Statistical methods used in fisheries work are often based on population-abundance methods, and more recently overlap with work more readily classed as environmental, such as water pollution monitoring. Line transect sampling methods and capture-recapture data are two of the many approaches that have been used, but much more biometric work could be done in this field.
4.2. Forestry
Various applications are used for forestry and agroforestry, including wood-volume estimation, sampling for specific properties, multispecies systems and their sustainability, inter-tree competition and its effect on plantation growth, thinning response functions, yields of forest stands subject to fire and insect risks, and forest tree breeding.
4.3. Statistical Ecology and Biodiversity
Statistical ecology and biodiversity have a long history as an integral part of biometry, and the relationship between the environment and ecology has seen the development of much useful methodology from this perspective. As societal concerns have changed, coupled with remote sensing information and computer technology, ecological research has also undergone appropriate advances. Sample survey data, as well as other methods like distance sampling and capture-recapture samp-ling, are now incorporated, along with intensive site-specific data and remote sensing image data. Many developments are expected in the foreseeable future, in areas including: automated survey design algorithms; advances in model-based inference from survey data; a common inferential framework for wildlife population assessment methods; improved methods for estimating population trends; improved models for conservation management; advanced spatiotemporal models of ecosystems; and better incorporation of model selection uncertainty into inference.
4.4. Morphometrics and Stereology
Morphometrics is the quantitative measurement of shape, and statistical study of shape change; allometry is the study of size and its consequences. These are classical areas of application of biometrical techniques. Such techniques include conventional multivariate statistical techniques, phylogenetic analyses, and distributional analyses, as well as themes from plane and solid geometry to support biological insights into features of many different organs and organisms. Applications are widespread, from analysis of dinosaur material to analysis of children’s growth patterns, such as height and craniofacial changes. Ecophenotypy is a relatively new area of study concerned with correlations between shape and environment.
Stereology is the science of inference about three-dimensional structures, based on two-dimensional sections or one-dimensional probes. In medical and biological research, microscopic slides are thin slices through specimens of material. They can give information about the volume of different types of tissue in the specimen, about the area of membranes, and about the length of capillaries. Various formulae and approaches have been derived. Sampling may be problematic: many stereological formulae depend on the orientation of the section, and taking sections that are random in direction may be impossible. A modeling approach from crystal growth studies, the Dirichlet tessellation of a point process, has been widely used to study plant ground-cover, and animals’ territories, in ecological applications. The related areas of image analysis and mathematical morphology are also ever-growing, with a newly-emerged application being the analysis of microarray images.
4.5. Bioassay and Toxicology
Bioassay refers to the process of evaluating the potency of a stimulus (a drug, hormone, radiation, environmental effect, toxicant) by analyzing the responses it produces in biological organisms (such as experimental animals, human volunteers, living tissues, bacteria). The familiar techniques of regression analysis and generalized linear models are used extensively. Experimental design for bioassay involves selection of dose levels, and allocation of living organisms to these levels. A number of sequential designs have been proposed, as information becomes available. Selection of dose levels for estimating the ED50, and extreme percent points (for example, ED90) can be based on Bayesian principles. When considering combinations of two or more drugs, more complex approaches have been developed to handle the problem of the drugs’ interaction.
Modern toxicology offers a wide range of sensitive tests for studying adverse health impacts, including tests for carcinogenicity, mutagenicity, and developmental toxicity. These three endpoints are of particular interest to biometricians, because of the challenging statistical problems arising from analysis of the experimental data. Continuing research has seen the development of biologically-based models that take account of pharmacokinetics, metabolism, and physiological systems to better characterize dose response, and to improve extrapolations from animals to humans. Toxicity may also be evaluated directly, such as assessment of pollutants in aquatic systems.
4.6. Infectious Disease Epidemiology
Infectious diseases have always been a major cause of death globally. For specified infectious diseases there exist international, national or local surveillance data in many countries today. The quality of such data varies greatly. Interpretation and comparisons are usually difficult, due to under-reporting, reporting delays, and diagnosis errors. Besides HIV/AIDS, data about measles and influenza have been the best studied. Many models, both deterministic and stochastic, have been invoked to explain the recorded patterns of measles and influenza outbreaks. These infectious diseases are subject to public health control actions, like vaccination or therapy, and so to be able to interpret reported incidence figures one also needs data about the quantity and quality of interventions. For detailed understanding of the transmission dynamics one needs to be able to analyze data at the household level. Often the parameters estimated from such data cannot be generalized to other sections of the population, however, so other information based on cross-sectional surveys is required to assess the epidemiological situation in the general population. Modeling seroprevalence data is possible, but not entirely straightforward, and further work on this would be beneficial.
The most important concept in epidemic theory is the average
number of secondary cases of an infectious disease that one case might generate
in a completely susceptible population. This is the parameter of fundamental
importance in infectious disease epidemiology known as the “basic reproduction
number,” R0. The higher the value of R0
then the more infectious is the disease. However, if R0<1,
transmission of the infection cannot be sustained and the infection will
eventually die out. Many assumptions need to be made in the various approaches
to estimating this number: for example, in the 1990s, an estimate of R0
of 3.8 was found for data on hepatitis A in
Evaluation of vaccines and vaccination programs is important, and a variety of relevant models have been developed to analyze vaccine trial data. The potential effects of vaccines have been widely discussed with respect to HIV. Some argue that prevention of HIV should not be the only endpoint for vaccine trials, and that a vaccine would also have great public health benefits if it reduced the progression rate of disease, or the infectivity of the virus. The same issues are also relevant to other infections, like malaria, and much more needs to be done in this area. In considering infectious diseases that are transmitted directly or indirectly by means of vectors (for instance, mosquitoes) from person to person, vaccination also protects the community. For a homogeneously mixing population and a random distribution of vaccines, the maximum non-immunized proportion required for the elimination of an infection equals the inverse of the basic reproduction number. It also has been shown that the concept of “herd immunity” is strictly only applicable for random mixing populations, and epidemics may occur in highly-vaccinated populations if the non-vaccinated are clustered. Because there is no possibility to conduct randomized, controlled trials with whole nations in order to compare alternative vaccination strategies, there is no substitute in the evaluation of vaccination programs but to use mathematical models to assess the likely effects of alternative interventions.
The HIV/AIDS epidemic has posed unique methodological problems. The incubation time (the interval between infection and onset of full-blown AIDS) is very variable, with a median of about 10 years, and it is still not clear if there is a subgroup of infected individuals who will never develop AIDS. The interpretation of epidemiological AIDS data is particularly difficult because of length-biased sampling, reporting delays, and under-reporting. There is a vast literature on the mathematical modeling and statistical analysis of HIV/AIDS data, and the interested reader is referred to the bibliography.
Besides development of complex mathematical modeling for the above-mentioned infectious diseases, there have been equally important developments in the modeling of influenza, polio, and smallpox, amongst other infectious diseases, and more recently for Foot-and-Mouth disease (FMD) and bovine spongiform encephalopathy (BSE) in animals, and Creutzfeldt-Jakob Disease (vCJD) in humans.
4.7. Genetics
It is useful to distinguish between the sub-areas of population genetics
and statistical genetics, although there are several areas of overlap.
Population genetics is concerned with the genetic variation seen in natural and
artificial populations. The various processes and properties that generate the
observed patterns of variationpopulation
structure, mating patterns, mutation, migration, genetic drift and natural
selection
have
been mathematically modeled, to determine the more important features. By
comparing the predictions of these models, the relative importance of the
different population genetic processes can be deduced for a particular genetic
system in the population of interest. Mathematical theories are also being used
to infer features of the evolutionary history of groups of organisms.
Statistical genetics is the development and application of statistical methods for genetic data. The genetic data may be allelic data from specific genes or from marker loci (like single nucleotide polymorphisms, SNPs), or phenotypic data from which inference about the genetic components of the phenotypes is to be made. Fields of application are various, including plant and animal breeding, and human genetics. In an experimental population individuals can be chosen for breeding, design is a major issue, and missing data are few. In human genetic studies, problems of missing data and sampling procedures dominate. Another major division between modeling approaches is between models to assess the extent of the genetic contribution to a trait only from trait data among relatives, and models that include the effects of specific genes or regions of the genome. Analysis at the genomic level is now falling under the new descriptor “bioinformatics.”
4.8. Bioinformatics and Genomics
The convergence of biological and information sciences will revolutionize global markets through the 21st century, transforming health care, pharmaceutical, and agricultural industries. Few areas of medicine and agriculture will be untouched by these advances in bioinformatics.
Bioinformatics is an emerging field of science, growing from the application of mathematics, statistics, and information technology to the study of very large biological, and particularly genetic, datasets. The increase in DNA data generation, in particular from the genome projects, has seen the development of many methods for the analysis of these massive datasets. Biologists need evaluation tools for these data, such as for the statistical assessment of similarity between two or more DNA or protein sequences, for finding genes in the genomes, and for estimating differences in gene expression in different tissues. Development of such tools requires biometrical modeling of biological systems. While this is currently a rapidly-growing field, it is also one where a greater incorporation of the underlying stochastic nature of the data, and less reliance on ad hoc techniques, would often be beneficial.
The comparison of DNA and protein sequences from different organisms (comparative genomics) is being used to help identify features of biological interest. Such information can be drawn on, for instance, to translate the genetic basis of disease susceptibility of a mouse mutant into the search for genes that affect an analogous human disease. An example is diabetes. Challenges in genomic epidemiology include the identification of the constellation of interacting genes that are principally responsible for conferring risk of common diseases. For example, with 30,000 human genes, there are 4,499,550,010,000 (4.5E12) possible three-gene interactions for any phenotype. Functional genomics encompasses protein and RNA structural prediction, along with the design, analysis, and interpretation of experiments using expression array data. The new array technologies produce the most comprehensive description of phenotype ever, with as many as 70,000 data points for a single sample. The biometrical challenges are exciting.
4.9. Public Health and Biomedicine
Biometrics’ impacts on epidemiology, and its resulting impact on public health and biomedicine, are fundamental, and so are difficult to overstate. Contemporary medicine is based on sound scientific methods, and the following section summarizes some biometrical approaches to clinical data.
Evidence-based medicinethe
integration of individual clinical expertise with a critical appraisal of the
best available clinical evidence from systematic research
is
a relatively young sub-discipline that is now being increasingly practiced.
Other evidence-based fields also are emerging. Since the randomized trial, and
especially the systematic review of several randomized trials, are much more
likely to inform clinicians, and so less likely to mislead them, than trials
where there is no randomization, randomized trials have become the “gold
standard” for evaluating whether a treatment does more good than harm. However,
evidence-based medicine is not restricted to such trials. The following
subsections outline the three major types of clinical studies.
4.9.1. The Case-Control Study
The sophisticated use and understanding of case-control studies is an outstanding methodological development in modern epidemiology. The central idea is the comparison of a group having the outcome of interest to a control group, with regard to one or more characteristics. The method can be traced back to 1843, and became popular in the 1920s, especially for the study of cancer. Following widespread criticism (mainly centered around claims of causality, and the smoking/lung cancer debate), J. Cornfield is credited with the start of the modern era in this approach in the 1950s, by showing that the relative risk (the exposure odds ratio) approximates the disease rate ratio, under appropriate sampling. The next major breakthrough was the elegant Mantel-Haenszel summary relative risk estimator, based on the stratification of data into a series of 2x2 tables for control (and visual evaluation) of confounding. A variety of likelihood methods have been developed over the past few decades, and today logistic regression is widely used. In the 1970s, methods to accommodate matching and nesting were introduced, while the 1980s saw more informative sampling schemes introduced, and development of modern statistical tools for model-validation and outlier-detection. Recently, semiparametric methods and visualization tools have become more widely used, along with computer-intensive approaches.
Despite all the technical advances, the fundamental problem of drawing causal inference from observational data remains. The limitations of case-control methodology are well known, including selection bias, measurement error, and confounding. More complex approaches continue to be developed to deal with these difficulties in an ever-more realistic manner. Currently, graphical models are being used to model causal reasoning based on directed, acyclic, graphs.
Case-control studies are not restricted to medical studies, although the methodological developments have been stimulated in this environment. Such studies have been used elsewhere, an example being habitat-association studies in wildlife research.
4.9.2. The Cohort Study
Cohort studies play a central role in epidemiological research and biometrical design, and analyses underpin their execution. Disease associations that are relatively strong can often be reliably studied using these techniques, particularly if there is sufficient knowledge of disease risk factors and exposures to control confounding. Cox regression and logistic regression methods play a central role in data analysis, as well as in study planning. There is an enormous literature on cohort study design, conduct, and analysis. The impact of covariate measurement error on these stages, as well as interpretation of cohort study results, is one of the least developed and potentially most important aspects of cohort study methodology. Further, analytic methods to reduce the sensitivity of results to measurement error influences are under development. Use of multi-population cohort studies to enhance exposure heterogeneity has the advantage that random effects modeling, say, can be invoked to partition the between-population and within-population components. Much of the retrievable information often arises from between population sources.
4.9.3. Clinical Trials
Based on the pioneering efforts of Bradford Hillwho
used the “new methodology” in the 1940s to help demonstrate the value of
streptomy in therapy for tuberculosis
clinical
trials incorporating the Fisherian principles of randomization and replication
today play a key role in the evaluation of all newly-proposed procedures.
Clinical trials design draws on the principles of statistical surveys and
experimental design, and usually there are multiple endpoints. To avoid bias,
analyses have been customized to accommodate patients who fail before treatment
can be completed, who withdraw from the trial, or who do not meet the protocol.
In particular, methodology has been developed for dealing with complex censored
survival data. Recent developments have also seen the application of
meta-analysis techniques. Although increasingly more randomized trials are
being conducted, there remain three main concerns: quality (in particular,
reaching high administrative quality across all trials research in all
countries); size (as most clinical trials are too small, and often follow-up is
too short); and relevance (ensuring that the right trials are being done, and
their results appropriately incorporated into improving clinical practice).
The above has predominantly concentrated on consideration of statistical methods in biometry. In this section the focus is on the more direct use of mathematics in biology. Biological complexity derives from the fact that biological systems are multifactorial, dynamic, and stochastic. Mathematics has had a major impact on three broad areas: cellular and molecular biology; organismal biology; and ecology and evolutionary biology.
The application of mathematics to cellular and molecular biology is so widespread that it is taken for granted. For example, the determination of the dynamic properties of cells and enzymes, expressed as enzyme kinetic measurements or receptor-ligand binding, is based on mathematical concepts that form the base of quantitative chemistry. The utility of the core tools of molecular biology was validated by mathematical analysis. Some examples are the quantitative estimates of viral titers, measurement of recombination and mutation rates, validation of radioactive decay measurements, and measurement of genome size and informational content based on DNA complexity. Today mathematics underpins methods to find genes, as well as in the shot-gun sequencing methods that have been used to obtain various genomes, including the human genome. In the area of structural biology, mathematics has made important contributions. This area lies at the interface of biology, mathematics, and physics, and sophisticated physical models have been developed to determine the structures of biologically-important macromolecules, their assembly into specialized particles and organelles, and at higher levels of organization. The grand challenges in this area relate to genomics and structural biology, including structural analysis, molecular dynamic simulation, and drug design.
Organismal (or organismic) biology deals with all aspects of the biology of individual plants and animals, including physiology, morphology, development, and behavior. Mathematicians have made many contributions, ranging from technological advances to developing theories of biological structure and function. In this area major challenges include the study of complex hierarchical biological systems, such as problems in neuroscience, immunology, genomic regulatory networks, and developmental biology. Another ongoing challenge is improving appreciation of the dynamic aspects of structure-functional relationships.
Ecology and evolutionary biology encompasses a broad range of levels of biological organization, from the organism through the population to communities and whole ecosystems, and a large range of spatial and temporal scales. For example, population biology deals with the basic aspects of ecological and evolutionary change, and includes a diverse range of topics. These include the construction of evolutionary trees from datasets, the interface of game theory and population genetics, the ecology and evolution of quantitative characters, molecular evolutionary dynamics, and human population genetics. Grand challenges are many: from global change, including biodiversity, to molecular evolution, including building bridges between population biology and molecular biology.
A number of basic mathematical issues are common to the above challenges:
1. How to incorporate variation among individual units in nonlinear systems.
2. How to incorporate interactions among phenomena that occur on a wide range of scales, of space, time, and organizational complexity.
3. Understanding the relation between pattern and process.
The uniqueness of biological systems shaped by evolutionary forces poses unique problems, and solving these is continuing to lead to the development of new mathematics.
Biometry is a vital and continually growing subject that complements and guides empirical research, elucidates mechanisms, and can provide model systems for biological study. Today we are seeing a move towards what is being called “in silico” biology, namely massive datasets, and complex analyses, for which large computational power is required. Biometrical methods underpin this approach. While we are seeing biologists’ dependence on mathematics and statistics increase, simultaneously we are starting to observe the development of new areas of mathematics and statistics, inspired by the questions posed by the “new biology.”
Computer power will continue to grow, leading to further development of biometrical methodology and more realistic modeling of biological systems. Today this is particularly pronounced in the “post-genome” information age. We are in the middle of this evolution, or what many refer to as a revolution.
Papers in most biologically based journals have a relatively short “half-life”: our knowledge of biological systems is recent, and so most of our biological theories are still rapidly evolving. By contrast, mathematical fact is immutable. Successful mathematical theories can have long lifetimes, and so papers in biometry and the parent disciplines of mathematics and statistics often have a much longer “half-life“ of twenty years or more. A recurrent problem has been the lag between advanced theory and current practice. Most biologists today have done at least an introductory course in statistics, but their understanding is too often insufficient to perform well-designed experiments, or to undertake effective analysis of their data. The development of research teams, with the biometrician playing a pivotal role, solves this problem.
Biometry today is a dynamic field that is relevant to many important major scientific and social issues that face us now. Biometrical methodology, some of which is outlined in the first three sections of this theme, has played a central role in the interpretation of experimental data in a wide range of biological and medical research, including the selected examples of biostatistics and biometrics topics listed below. Biometry as an enabling discipline has been fuelled by practical problems in public health, forestry, genomics, ecology, and environmental contamination, to name a few. A major requirement in biometry today is to train more biometricians to meet current demands.
Click Here To View The Related Chapters
|
Accelerated failure time model |
:Let T0 be the survival time, under control conditions, from some origin to the occurrence of an event of interest, and suppose that application of a treatment modifies the survival time to T=T0/θ for some scaling parameter θ. In its simplest form, the accelerated failure time model involves proportional adjustment of the time scale, so the median (X percentile) survival time under the treatment is 1/θ times the median (X percentile) under the control. |
|
AIDS |
:Acquired Immune Deficiency Syndrome; certain conditions of poor health in HIV-infected individuals qualify them for diagnosis of AIDS. |
|
Allele |
:Different states (nucleotide sequences) of a gene. |
|
Analysis of variance |
:An analysis based on separating sums of squares into components associated with defined sources of variation used as criteria of classification for the observations. |
|
Ancillary statistic |
:A statistic that does not depend on the parameter θ of the distribution. |
|
Asymptotic |
:In the limit. |
|
Autocovariance |
:The covariance between (usually) time points at fixed times apart. |
|
Autoregression |
:The generation of a series of observations whereby the value of each observation is partly dependent on those that have immediately preceded it. |
|
Back-calculation |
:Also known as back-projection, this, for example, estimates past infection rates of an epidemic by working backward from observed disease incidence, using knowledge of the incubation period between infection and disease. |
|
Bayesian methods |
:These are based on the result that the posterior probability of B, given A, is equal to the prior probability of B times the likelihood of A, conditional on B divided by the probability of A. |
|
Bias |
:Systematic error, as distinct from random error that may distort on any one occasion but balances out on average. |
|
Binomial distribution |
:If an event has
a probability p of occurring in a trial, the probability of r
events occurring in n independent trials is n!prqn |
|
Biomathematics model |
:The term is used synonymously with biometry, to stress that the subject also involves non-statistical methods. |
|
Biplot |
:A scattergram of a two-dimensional data array, in which both the rows and columns are represented by points. |
|
Birth and death process |
:A stochastic process that attempts to describe the growth and decay of a population, whose members may die or give birth to new individuals. |
|
Bootstrap |
:Bootstrap methods are procedures for the empirical estimation of sampling distributions and their properties. They are used in situations where limited data are available, and traditional analyses are either difficult or unreliable. |
|
Branching process |
:A stochastic process describing the growth of a population in which the individual members may have offspring. |
|
Causal relation |
:A deterministic or stochastic relation between a cause and the corresponding response. |
|
Classification |
:A term often used as synonymous with discriminant analysis or cluster analysis. |
|
Cluster analysis |
:A general approach to multivariate problems, in which the aim is to see whether the units fall into groups or clusters. |
|
Conditional statistic |
:A statistic whose distribution depends upon some quantity that is held constant; the quantity in question is usually itself some function of the variables entering into the statistic. |
|
Confounding |
:When certain comparisons can be made only for treatments in combination and not for separate treatments, those treatment effects are said to be confounded. Confounding can be a deliberate feature of the design, but may arise from inadvertent imperfections in a study. |
|
Correlation |
:A correlation coefficient
is a measure of the interdependence between two variates. It is usually a
number between |
|
Correspondence analysis |
:An analogue for contingency table analysis of Principle component analysis for frequency data, and a method to solve the problem of ordination. |
|
Covariance |
:The first product moment of two variables about their mean values. |
|
Differential equation |
:An equation that contains, in addition to the independent variables and one or more unknown functions, derivatives of those functions. |
|
Dirichlet tessellation |
:Associated with any realization of a point process in space is a system of polygons in two dimensions, or polyhedra in three dimensions, containing those points closest to each point of the process. This is known as a Dirichlet tessellation. |
|
Discriminant analysis |
:Determination of a rule to allocate individuals to their correct population, with minimum probability of misclassification. |
|
ED50 (ED90) |
:Effective dose 50% (90%). |
|
EM algorithm |
:A general computational method for calculating maximum likelihood estimates with incomplete data. Each iteration has two steps, an E-step for computing the estimation of the missing data, and an M-step for computing the maximum likelihood estimates of the parameters assuming complete data. |
|
Factorial design |
:Used in an experiment for testing the effect of a set of treatment factors. |
|
Fourier analysis |
:The theory of representing functions of a variable t as the sum of a series of sine and cosine terms of type ajcos(2πj/λj), j=0,1, … |
|
Gaussian |
:See, Normal distribution. |
|
Generalised estimating equations (GEEs) |
:The GEE approach makes weaker distributional assumptions than are required for a fully parametric, likelihood-based approach, while maintaining the properties of consistency and asymptotic normality of the parameter estimates. |
|
Genomics |
:The study of the relationships between gene structure and biological function in organisms. |
|
Gibbs sampling |
:A special case of the Metropolis-Hastings algorithm, used in MCMC for constructing the relevant Markov chain. |
|
HIV |
:Human Immunodeficiency Virus. |
|
Kalman Filter |
:An iterative technique of dynamic linear modeling, used mainly for estimating the parameters of autoregressive moving-average time series models with Normal residuals. |
|
Kernel estimators |
:These are convolutions of a smooth function with a rough empirical function, in such a way as to produce a smooth functional estimator. |
|
Latin square |
:A scheme for arranging v numbers in v rows and v columns, so that each number appears exactly once in each row and column. |
|
Likelihood |
:The probability or density function considered as a function of the parameter for a given realization of the random function. |
|
Logistic regression |
:A model often considered suitable for prediction for binary (dichotomous) data; it is linear in the logit (logarithm of the odds) as a function of the explanatory variables. |
|
Log-linear model |
:A model often considered suitable for prediction from data of contingency table form; it is linear in the logarithms of the theoretical frequencies of the contingency table. |
|
Markov process |
:A stochastic process such that the conditional probability distribution for the state at any future time, given the current state, is unaffected by knowledge of the past history of the system. |
|
Maximum likelihood |
:The value of the unknown parameters that maximize the likelihood function. |
|
Model |
:A formalised expression of a theory that is regarded as having generated the observed data. |
|
MCMC |
:Markov chain |
|
Multidimensional scaling |
:The derivation of a structured multidimensional scale from empirical data. |
|
:Also known as
Gaussian distribution. The continuous frequency distribution represented by dF=exp( |
|
|
Ordination |
:Reduction of dimensionality of multivariate data to an “ordering,” originally along a single axis, but also to two- or three-dimensional representations. |
|
Outlier(s) |
:Observation(s) so far separated in value from the remainder that the question arises whether they are from a different population, or if the sampling technique is at fault. |
|
p-value |
:Often used in biology to denote probability level (α for significance level); the probability if the null hypothesis is true of a test statistic value as extreme as, or more extreme than, the value observed. |
|
Parameter |
:An unknown quantity that may vary over a certain set of values. |
|
Pharmacokinetics |
:The study of the bodily absorption, distribution, metabolism, and excretion of drugs. |
|
Phylogenetic analysis |
:Determination of the evolutionary relationship, usually of tree-like structure. |
|
Point process |
:A statistical process concerned with the occurrence of events at points of time determined by some chance mechanism. |
|
Poisson distribution |
:A discontinuous
distribution with the probability of the random variable having value r
being λre |
|
Power |
:In general, the power of a statistical test of some hypothesis is the probability that the null hypothesis is rejected when a simple alternative is correct. |
|
Prediction interval |
:The interval between the upper and lower limits attached to a predicted value to show, on a probability basis, its range of error. |
|
Principal components |
:The variables obtained by a linear transformation of a multivariate set of data, so that the newly derived variables are uncorrelated and each accounts in turn for as much of the variation as possible. |
|
Probability |
:A basic concept that may be taken either as undefinable, in expressing a “degree of belief,” or as the limiting frequency in an infinite random series. Both approaches lead to the same calculus of probabilities. |
|
Proportional hazards model |
:A model assuming that the factors affecting survival have an additive effect on the log hazard function. |
|
Proteomics |
:The study and analysis of protein structure and function. |
|
Randomization |
:Assignment of treatments to experimental units, with the aim of guaranteeing the randomness of the sample. |
|
Randomized trial |
:Randomization of subjects to different treatments, or levels of treatment. |
|
Replication |
:The performance of an experiment, survey, or study more than once, so as to increase precision and obtain a better estimation of sampling error. |
|
Robustness |
:A statistical procedure is said to be robust if it is not very sensitive to departures from the assumption(s) upon which it is based. |
|
Sequential analysis |
:The analysis of data derived by a sequential method of sampling. In sequential sampling, the units are drawn in order, and the results of the drawing at any stage decide whether or not the sampling is to continue. Analyses of such data (including data from early termination of a study) differ from fixed sample methods. |
|
Spectral analysis |
:The spectral representation of the autocovariance function of a stationary process. |
|
Stationary process |
:A stochastic process {xt} is said to be strictly stationary if the multivariate distribution of {x} with successive subscripts t1+h, t2+h, …, tn+h is independent of h for any set of parameter values t1+h, t2+h, …, tn+h, t1, t2, …, tn. |
|
Stochastic model |
:A model that incorporates some stochastic (random) elements. |
|
Sufficiency |
:A technical term for a statistical property of an estimator, defined by R. A. Fisher (1921). |
|
Variable |
:A quantity that may take any one of a specified set of values. |
|
Variance |
:The second moment taken about the mean; the mean of the squares of variations from the arithmetic mean. |
Anderson,
R. M.; May, R. M. 1991. Infectious Diseases of Humans: Dynamics and Control.
Armitage,
P.;
Armitage,
P.; David, H. A. (eds.) 1996. Advances in Biometry.
Biometrics. http://stat.tamu.edu/Biometrics/ [Scientific journal of the International Biometric Society, containing the latest biometrical research.]
Bookstein,
F. L. 1997. Morphometric Tools for Landmark Data: Geometry and Biology.
Bürger,
R. 2000. The Mathematical Theory of Selection, Recombination, and Mutation.
Carlin,
B. P.; Louis, T. A. 2000. Bayes and Empirical Bayes Methods for Data
Analysis. 2nd edn.
Diggle,
P. 1990. Time Series: A Biometrical Introduction.
Ewens, W.
J.; Grant, G. R. 2001. Statistical Methods in Bioinformatics: An
Introduction.
Johnson,
N. L., Kotz, S. (eds.) 1988. Encyclopedia of Statistical Sciences.
Journal of Agricultural, Biological and Environmental Statistics (JABES). http://www.tibs.org/jabes/ index.html. [Major scientific journal.]
Lange,
N.; Ryan, L.; Billard, L.; Brillinger, D.; Conquest, L.; Greenhouse, J. 1994. Case
Studies in Biometry.
Levin, S. A. (ed.). Mathematics and Biology: The Interface Challenges and Opportunities. http://www.bis.med.jhmi.edu/ Dan/mathbio/T.html. [Report on opportunities at the interface between biology and mathematics.]
Maindonald, J. 2000. The Design of Research Studies: A Statistical Perspective. http://www.anu.edu.au/graduate/pubs/ occasional-papers/gs00_2.pdf. [Excellent notes addressing broad planning principles that apply to many research areas; source of references to various biometrical areas, including experimental design and evidence-based medicine.]
Manly, B.
F. 1991. Randomization and
McCullagh,
P.; Nelder, J. A. 1989. Generalized Linear Models. 2nd edn.
Methodology Group, NHS R&D Health Technology Assessment Programme. http://www.hta.nhsweb.nhs .uk/. [Downloadable monographs and review papers on health technology.]
Rasch,
D.; Tiku, M. L.; Sumpf, D. 1994. Elsevier’s Dictionary of Biometry.
Sokal, R.
R.; Rohlf, F. J. 1995. Biometry: The Principles and Practice of Statistics
in Biological Research. 3rd edn.
Statistics in Medicine. http://www.interscience.wiley. com/. [Major scientific journal covering biostatistics.]
Statistical Methods in Medical Research. http:// www.smmrjournal.com. [Major scientific journal covering biostatistics.]
Taylor,
H. M.; Karlin, S. 1998. An Introduction to Stochastic Modeling, 3rd edn.
Welsh, A.
H. 1996. Aspects of Statistical Inference.
Sue
Wilson is Professor and Head, Statistical Science Program, Centre for
Mathematics and its Applications, School of Mathematical Sciences, and
Co-Director, Centre for Bioinformation Science (joint with John Curtin School
of Medical Research), at the Australian National University (ANU). She obtained
her B.Sc. from the
Sue has over 150 publications in biometry and applied statistics, with a particular emphasis on statistical genetics/genomics. These papers have arisen from her extensive consulting experience in the biological, social, and medical sciences, leading to statistical modeling developments to answer substantive research questions in these disciplines. She is currently involved in the establishment of a bioinformatics research facility at ANU.
Sue is an
elected member of the International Statistical Institute, a Fellow of the
American Statistical Association and a Fellow of the Institute of Mathematical
Statistics (IMS). She was President, International Biometric Society, 19992000
(Vice President, IBS, 1998, 2001). Currently she is Associate Editor, Annals
of Human Genetics; Associate Editor, Computational Statistics and Data
Analysis; Member, Editorial Board, Statistical Methods in Medical
Research; Member, Conference Advisory Committee, IBS; Member, IMS Committee
on Memorials; Member, IMS Nominations Committee; Member, Editorial Committee
for the 6th edition of the ISI’s Dictionary of Statistical Terms.
|
To cite this chapter |
| ©UNESCO-EOLSS | Encyclopedia of Life Support Systems |