3.3: Selecting an appropriate method
- Page ID
- 116750
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)The method of choice for assessing food or nutrient intakes depends primarily on the objectives of the study. No method is devoid of random or systematic errors (Chapter 5), or prevents alterations in the food habits of the respondents. However, statistical techniques are now available which are designed to mitigate the impact of measurement errors on study results. See Freedman et al. (2011), Kirkpatrick et al. (2018), and Chapter 5. The US National Cancer Institute has developed a Dietary Assessment Primer (2020), a web resource to aid researchers choose the best available dietary assessment approach to achieve their research objective. Readers are advised to consult the primer before selecting a self-report dietary instrument.
Box 3.2 provides guidance on the most appropriate methods for assessing food or nutrient intakes in relation to four possible levels of objectives.
Level One: Mean nutrient intake of a group
- Preferred approach A single 24h recall, or single weighed or estimated food record, with large number of subjects and adequate representation of all days of the week
- Preferred approach Replicate observations on each individual or a subsample using 24h recalls or weighed or estimated one day food records with an adequate representation of all days of the week
- Preferred approach Multiple replicates of 24h recalls or food records or a semiquantitative food frequency questionnaire
- Preferred approach Even larger number of recalls or records for each individual. Alternatively, a semiquantitative food frequency questionnaire or a dietary history can be used.
Note that the number and selection of replicate 24h recalls, or weighed or estimated one day food records required to obtain level two, level three, or level four data, depends on the day-to-day variation within one individual (i.e., within-person variation ) of the nutrient of interest (Chapter 6). This variation depends on the nutrient, the study population, and the seasonal variations of intake. Generally, for nutrients found in high concentrations in only a few foods, such as vitamins A and D and cholesterol, the number of replicates needed is greater than for those found in a wide range of foods (e.g., protein). Nonconsecutive days should be selected for the replicates when possible, to enhance the statistical power of the information: day-to-day correlations between intakes often occur when food intake data are collected over consecutive days. The length of time needed between the observation days also depends on the nutrient (IOM, 2000; a 3–10d interval after the previous recall is recommended (Tooze, 2020).
Additional factors that should be considered when choosing a method for assessing the food consumption of individuals are the characteristics of the individuals within the study population, the respondent burden of the method, and the available resources. For instance, certain methods are unsuitable for elderly subjects with poor memories, for busy mothers with young children, or for illiterate individuals. Other methods require highly trained personnel and specialized laboratory and computing facilities, which may not be available. Generally, the more accurate methods are associated with higher costs, greater respondent burden, and lower response rates. Unfortunately, compromises often have to be made between the collection of precise data on usual nutrient intakes of individuals and a high response rate.
3.3.1 Determining the mean nutrient intake of a group: level one
Level one is the easiest objective to achieve and can be met by measuring the food intake of each subject in the group using a single 24h recall or a one day food record, provided the individuals are representative of the study population and all the days of the week are proportionately represented in the final sample. Data on mean usual nutrient intakes of a group can be used for international comparisons across countries of the relationship of nutrient intakes to health and disease. However, this method based on a single intake day per subject should never be used for reporting the distribution of intakes (i.e., as percentiles of intake) (Tooze, 2020). To calculate n, the number of subjects required in the group, an estimate of the between-person variation for the nutrient of interest is needed. This is usually obtained from the literature but may be determined during a pilot study.
As an example, assume the expected mean iron intake obtained from the literature is 10mg/d, with an anticipated standard deviation (s) of 3mg/d. Also assume that we want to be 95% confident that the true mean lies between 9.2 to 10.8mg/d (i.e., the confidence interval has limits which are 0.8mg/d on either side of the mean). A 95% confidence interval is calculated approximately as the mean ±2 × e, where e is the standard error of the mean — a measure of the precision of the estimated mean. Hence , the required e = 0.8/2 = 0.4. We can use the following formula to calculate n, the desired group size.
\[n=s_b^2 / e^2\nonumber\]
where \(s_b^2\) is the between-person variance of the nutrient of interest, and e is the desired standard error — a measure of precision required for the estimate of the mean intake of the nutrient of interest. Hence
\[n=3^2 / 0.4^2=56.25\nonumber\]
and so 56 subjects are required. Alternatively, if we wanted to be 99% confident that the true mean lies between 9.2 and 10.8mg/d, then more subjects must be studied. A 99% confidence interval is calculated approximately as the mean ±3 × e. Hence the required e = 0.8/3 = 0.27 Therefore
\[n=3^2 / 0.27^2=123\nonumber\]
and 123 subjects are required.
Clearly, the size of the group (n) necessary to characterize the group mean usual nutrient intake depends on the degree of precision required: more individuals must be studied to achieve higher precision. This calculation of the required group size should be repeated for each of the nutrients of interest, and the largest n (i.e., the worst case) should be used if possible.
If the study objective is to demonstrate a significant difference in the mean intakes of two groups, or a significant change in the mean intakes, based on unpaired or paired data, then alternative formulae must be applied; details are given in (Gibson and Ferguson, 2008).
3.3.2 Calculating the population percentage “at risk": level two
To determine the percentage of the population “at risk” of inadequate nutrient intakes, an estimate of the distribution of usual intakes of the individuals is required. This, in turn, requires that the food consumption of individuals be measured over more than one day. Hence, repeated 24h recalls, or replicate weighed or estimated one day food records are the methods of choice, again ensuring that all days of the week are proportionately represented in the final sample. Often, it is not feasible to carry out repeated observations on all the individuals, as in the case of a national dietary survey, and the recalls or records are repeated on a subsample of the individuals only.
To achieve a level two objective, at least two independent measurements of food intake should be obtained on at least a representative random subsample of individuals in the survey. The U.S. Food and Nutrition Board (IOM, 2000) recommends that the replicate measurements should be independent and made on non-consecutive days 3–10d after the first recall or record. However, if the data can be collected only on consecutive days, then three daily measurements should be used. The random subsample should consist of at least 50 individuals per demographic group. Where possible, the demographic groups to be sampled should be defined according to the sex and life-stage groups used for the chosen Nutrient Reference Values (Deitchler et al., 2020),
Note that it is more important to have a minimum number of replicate observations in the subsample than a minimum proportion of replicate observations. Once the required number of replicate observations have been obtained, methods can be applied to correct for the measurement error associated with both day-to-day variability in intakes for a single person (termed within-person variation ) and the variation in intakes between individuals (termed between-person variation ) (see Chapter 6).
Four statistical methods are now available for estimating the distribution of usual intake. For more details, the reader is referred to (Tooze, 2020). Of the methods, the first was outlined by the National Research Council and later refined by Iowa State University (ISU) (Nusser et al., 1996) with software developed to use the ISU method. Note, this method does allow for a random sub-sample of repeated 24h recalls or food records, and has the capacity to adjust for season, day of the week, and/or sequence effects, but is not recommended for general use with episodically consumed foods, food groups, or nutrients (Tooze, 2020). The adjustment process provides estimates of the usual nutrient intakes for each specified age and gender-specific subgroup. An example comparing the adjusted distributions of usual zinc intakes (using the refined NRC approach of (Nusser et al., 1996) with the observed zinc intakes for New Zealand adult females aged 19–50y is shown in Figure 3.6.

Figure 3.6 Estimates of usual intake distribution for zinc for New Zealand adults obtained from 24-hout recall data and adjusted with replicate intake data using the refined NRC method. The y-axis (frequency of intake) shows the likelihood of each level of intake in the population. EAR, Estimated Average Requirement. Modified from (Gibson et al., 2003). Nutrition today, 38(2), 63–70.
The adjustment process used yields a distribution with reduced variability (sometimes referred to as “shrinking”) because the within person variation has been removed, while preserving the shape of the original observed distribution (Gibson et al., 2003).
Later, the National Cancer Institute (NCI) (Tooze, 2020) developed a method which also adjusts for the same covariates (season, day of the week, and/or sequence), but also has the capability of incorporating covariates to identify estimates for sub-populations in the survey. In addition, the NCI method, unlike the ISU, can be used with episodically consumed foods, food groups, or nutrients.
The European Food Consumption Validation Project have developed the Multiple Source Method(Haubrock et al., 2011) (MSM) that can also be used for episodically consumed foods, food groups, or nutrients, although caution must be used when employing the program with models containing covariates. The Dutch National Food Consumption Survey 2007-2010 have developed a program called the Statistical Program to Assess Dietary Exposure (SPADE) , but this program currently requires 24h recalls on at least two-days for all respondents in the sample and is not recommended for use with episodically consumed foods, food groups, or nutrients.
All the methods described briefly above use varying approaches to yield an adjusted distribution of “usual" nutrient intakes that can then be used to predict the proportion of the population at risk of nutrient inadequacy using either the full probability approach, or the Estimated Average Requirement (EAR) cutpoint method; details are given in Chapter 8b. However, the choice of which “usual intake method” method to use depends on whether foods, food groups, or nutrients are episodically consumed, whether the probability of consumption and the consumption-day amount are correlated or not, and whether replicates of the 24h recalls or food records are available on all the population groups to be surveyed or only on a sub-sample (Souverein et al., 2011).
| Variable | Prevalence (%) "True" Observed |
No. of repeated measurements needed |
|
|---|---|---|---|
| Cholesterol > 300mg |
15 | 37 | 39 |
| Calcium > 800mg |
12 | 21 | 9 |
Figure 3.6 shows that, in the example, adjusting the distribution significantly using the ISU method reduces the proportion of individuals considered to have intakes below the EAR for zinc. Within-person variation can also have a significant effect on estimates of the prevalence of abnormally high nutrient intakes. Table 3.7 shows data from NHANES II. The large differences between the observed prevalence and the calculated “true” prevalence represent the effect of removing the within-person variation by calculation. In this case, very large numbers of repeated measurements on each individual are required to reduce the observed prevalence of abnormally high intakes of cholesterol and calcium to within 5% of the true prevalence (Sempos et al., 1999). Comparable data are essential for national food policy development and food fortification planning. Food patterns associated with inadequate nutrient intakes can also be identified using this approach, enabling food assistance programs to be designed and improvements in nutrition education made.
More recently, a new method has been developed based on single-day dietary data which can be used to estimate population distributions of usual intake of nearly-daily consumed foods and nutrients, provided a suitable external within-person to between-person variance is available (Luo et al., 2019). Nevertheless, researchers are urged to collect replicate data where possible.
3.3.3 Ranking individuals by food or nutrient intake: level three
When the study objective is at level three and involves ranking individuals within a group, often for the purpose of linking dietary intakes with risk of chronic disease, the preferred approach is to obtain multiple observations on each individual. The number of days required to achieve the level three objective can be calculated from the ratio of the within- to the between-person variation in nutrient intakes (often termed the “variance ratio"); for more details see Chapter 6 . Sometimes an estimate of the variance ratio can be obtained from the literature, again preferably from an earlier study on a comparable group. Alternatively, a pilot study may be necessary to obtain this information.
Several authors have developed equations for calculating the number of replicate days required to meet level three objectives (Black et al., 1983; Basiotis et al., 1987; Nelson et al., 1989). Black et al. (1983) suggest using the following formula for the number of days (n) of diet records needed:
\[n=\left(r^2 /\left(1-r^2\right)\right) \times\left(s_w^2 / s_b^2\right)\nonumber\]
In this equation, r is the unobservable correlation between the observed and true mean intakes of individuals over the period of observation, and \(s_w^2\) and \(s_b^2\) are the observed within‑ and between-person variances, respectively. This equation should be used in association with Table 3.8 which shows the proportion of individuals correctly and incorrectly classified in the extreme fractions for different values of the corrlation coefficient between the observed and true intakes (r). The value of r chosen will depend on the degree of misclassification that the investigator is prepared to accept (Table 3.8).
| Correctly and incorrectly classified into extreme fraction |
||||
|---|---|---|---|---|
| r | Thirds | Fourths | Fifths | |
| 0.75 | a b |
0.69 0.049 |
0.63 0.013 |
0.59 0.004 |
| 0.80 | a b |
0.72 0.033 |
0.68 0.006 |
0.65 0.002 |
| 0.85 | a b |
0.76 0.018 |
0.72 0.002 |
0.69 <0.001 |
| 0.90 | a b |
0.80 0.006 |
0.77 <0.001 |
0.75 <0.001 |
| 0.95 | a b |
0.86 <0.001 |
0.84 <0.001 |
0.83 <0.001 |
As an example, assume that the investigator requires that when the individuals are divided into terciles, fewer than 5% (< 0.05) of the individuals are grossly misclassified into the opposite tercile. This will require an r value of 0.75 (Table 3.8). Assuming \(s_w^2 / s_b^2=1.7\), then the number of days (n)
\[\begin{gathered}
n=\left(r^2 /\left(1-r^2\right)\right) \times 1.7 \\
n=0.75^2 /\left(1-0.75^2\right) \times 1.7 \\
n=3 \text { days }
\end{gathered}\nonumber\]
The number of days needed to generate a given r increases as the chosen r increases. If the size of the within-person variation \((s_w^2)\) in nutrient intake is small compared with the size of the between-person \((s_b^2)\) variation , then fewer replicate days are needed to meet level three objectives.
The U.S. subcommittee on criteria for dietary evaluation (NRC, 1986) recommended using independent days for replicating the measurements of one day nutrient intakes to reduce any effect of autocorrelation between intakes on adjacent days.
An alternative approach to achieving level three objectives is to use a semi-quantitative food frequency questionnaire. This approach is often used in epidemiological investigations to study associations between intakes and risk of disease and does not require a measurement of absolute nutrient intakes. Although this approach is much simpler, involving only a single interview with each subject, it is difficult to quantify the errors involved and to separate the effects of within- and between-person variance.
3.3.3 Determining usual intakes of nutrients of individuals: level four
Reliable estimates of usual food or nutrient intakes of individuals that can be used with confidence to meet a level four objective, involving correlation or regression analysis with individual biochemical measures, are the most difficult to obtain. Large numbers of measurement days for each individual are required using 24h recalls or estimated or weighed food records.
An estimate of the within-person variation for each nutrient of interest should be obtained from the literature, preferably from an earlier study on a comparable group or a pilot study, as noted earlier. This estimate may be expressed as the variance, \(s_w^2\); standard deviation, \(s_w\); or as the coefficient of variation \((CV_w)\) expressed as a percentage:
\[C V_w=s_w /(\text { mean intake }) \times 100 \%\nonumber\]
This estimate can be used in the following equation to determine the number of days required per individual to estimate an individual's nutrient intake to within 20% of their true mean 95% of the time (Beaton et al., 1979).
\[n=\left(Z_\alpha C V_w / D_0\right)^2\nonumber\]
where n = the number of days needed per individual, Zα = the normal deviate for the percentage of times the measured value should be within a specified limit (i.e., 1.96 in the example below), CVw = the within-person coefficient of variation (as a percentage), and D0 = the specified limit (as a percentage of long-term true usual intake) (i.e., 20% in the example given below).
The following example calculates the number of days required to estimate a Malawian woman's zinc intake using 24h recalls to within 20% of the true mean, 95% of the time. In this example, the CVw (i.e., 34%) for zinc intakes on Malawian women via 24h recalls is taken from the literature (IZiNCG, 2004). Thus if Zα = 1.96 and CVw = 34%. then:
\[n=(1.96 \times 34 \% / 20 \%)^2=11 d\nonumber\]
If a pilot study is undertaken in which replicate 24h recalls are conducted, then the actual CVw for each nutrient of interest can be calculated. In this way, the estimate of the number of days required to measure the usual intake of each of the nutrients of interest in an individual, with a required degree of precision, can be defined. In general, considerably more days are required to obtain reliable estimates of intakes of individuals to meet the level four objective, compared with level three (i.e., relative ranking of subjects into groups) (Palaniappan et al., 2003).
Sometimes, dietary histories or semi-quantitative food frequency questionnaires are used to obtain this level four data on usual nutrient intakes for correlation with biomarkers (Jacques et al., 1993). Some investigators emphasize, however, that the accuracy of a semi-quantitative food frequency questionnaire is only equivalent to two to three repeat 24h recalls (Sempos et al., 1999).
In some experimentally controlled studies such as balance studies, information on the actual nutrient intakes of an individual over a finite time period are required. For such data, weighed food records (Section3. 1.4), completed for the duration of the study period, are the recommended method. Nutrient intakes can then be calculated using food composition data. Alternatively, duplicate meals can be collected throughout the period for later analysis; details are given in Chapter 4.
3.3.4 Combining dietary instruments
Recognition of the limitation of even two 24h recalls to adequately measure, at the individual level, usual intake of foods or nutrients that are not consumed daily (i.e., episodically consumed foods) has led to the use of more than one dietary instrument in a given study. In this way, the advantages of each method can be exploited and their weaknesses minimized. Consequently, in large scale epidemiological studies, a common approach is the administration of a food frequency questionnaire to all respondents, and 24h recalls or food records in a subsample. The food frequency questionnaire provides information on foods that are consumed less frequently (often, “in the past year") not given by the 24h recalls while the 24h recalls yield richer detail by attempting to correctly quantify portion size for each eating occasion (Subar et al., 2006). Alternatively, a food frequency questionnaire could be used as a supplement to one or more quantitative 24h recalls administered to the full sample, providing richer detail and less systematic bias. This approach appears to have the most added value when the research question of interest is to estimate diet-health relationships, especially when the estimation of usual dietary intake at the individual level involves a food, food group, or nutrient that is consumed episodically (Tooze, 2020) Statistical methods have been developed to combine the dietary data from repeated 24h recalls and food frequency questionnaires. Conceptually, these statistical methods presume that the usual food intake of an individual equals the probability of consuming a food on a given day (queried as frequency of usual intake over a specified time period), multiplied by the average amount of intake of that food on a typical consumption day. Repeated 24h recalls from the same individual yield information on consumption probability and amount, while the food frequency questionnaire provides information on intake frequency of rarely consumed foods and usual intakes of nutrients when portion size information is available (Conrad and Nothlings, 2017). The reported food frequencies and nutrient intakes (if available) can be used as covariates in both steps of the NCI method (Tooze, 2020) and with MSM (Haubrock et al., 2011) to enhance the estimation of usual intakes from the 24h recall data, especially from foods that are not consumed every day. Note, however, that in low income countries where context-specific food frequency questionnaires are often not available or inappropriate in the study setting, it may be preferable to increase the number of repeat 24h recalls per respondent rather than embarking on the development, validation, and collection of dietary data using a food frequency questionnaire in addition to using the 24h recall method. For more details, the reader is advised to consult (Tooze, 2020). at: Intake.org
3.3.5 Surveys of individual food consumption at the global level
FAO/WHO Global Individual Food consumption data Tool (FAO/WHO GIFT) is a global database of surveys of food consumption at the individual-level which are collected at the national and subnational level throughout the world. Only data collected from 1980 by quantitative methods such as 24h recalls, food records (weighed or estimated), 12hr recalls, and direct food weighing are considered. Portion sizes of all foods and beverages consumed by each survey participant must be assessed, including water if possible. For low‑ or middle-income countries, surveys comprising at least 100 subjects (with no evidence of strong selection bias) are included, whereas for high-income countries, surveys that are nationally representative are an additional criterion. The dietary data generated from FAO/WHO GIFT can be used to inform agricultural, nutrition, food safety and environmental policies and programs. A summary of 218 dietary surveys performed in low‑and middle-income countries from 1980 to 2019 and included in the FAO/WHO GIFT inventory is available in de Quadros et al. (2022).
FAO/WHO GIFT apply several criteria to validate the dietary databases included. For example, for all foods and drinks consumed by each survey participant on each survey day, a complete description, the amount consumed, and energy and nutrient values for each item must be included in the datasets. Additional compulsory variables are the age and sex of each subject, geographical location (country, and region(s) if available), type of area (e.g., rural, urban), and the number of survey days recorded per subject or for a subset of subjects. Inclusion of variables based on anthropometric data such as weight and height and the physiological status of both women (e.g., pregnant, lactation) and infants (i.e., breastfeeding status) are also recommended.
Recipes should be disaggregated whenever possible to provide data on the quantity of each separate ingredient, cooking method used, total amount prepared, number of people served from the recipe, and the quantity consumed by the survey participant (or served and leftover). Only with this information can each ingredient be attributed to their appropriate food group for the calculation of the FAO/WHO GIFT indicators and summary statistics.
The food composition values used are provided by the data providers. However, all individual quantitative dietary datasets shared through FAO/WHO GIFT are coded with the FoodEx2 system, a comprehensive and flexible food classification and description system. Eligible dietary databases are mapped manually by the investigator with FoodEx2 codes for food groups (n=24) and food subgroups based on the descriptions provided by FAO/WHO GIFT. FoodEx2 was first developed by the European Food Safety Authority (EFSA) and was later scaled up to the global level in collaboration with FAO and the World Health Organization (WHO). The use of this common food classification and description system among dietary surveys from different countries contributes to the global harmonization of dietary data. Prior to the analysis and formatting by FAO/WHO GIFT, all eligible datasets are screened for potential errors, missing values and outliers. For further details see FAO/WHO GIFT (Methods)
Provided the eligible dietary datasets have been coded using FoodEx2, the FAO/WHO GIFT platform can compute ready-to-use indicators and summary statistics based on individual food consumption data in the areas of food consumption, food safety and nutrition. For example, of the indicators, one reflecting dietary diversity at the population level (a key component of diet quality) entitled the Minimum Dietary Diversity for Women(MDD-W) is computed. This is a food group-based indicator that estimates the proportion of non-pregnant women of reproductive age (15‑49y) who consumed at least five out of ten defined food groups over the previous 24h.
The summary statistics generated by FAO/WHO GIFT comprise the estimated usual intakes of selected nutrients for pre-defined population groups by sex and age (except children less than aged 12 months) provided the datasets contain multiple non-consecutive days of 24h recalls / food records for at least a subset of 50 individuals. The statistical program chosen to adjust the dietary data for day-to-day variation is the Statistical Program to Assess Dietary Exposure (SPADE) by Dekkers et al. (2014). In addition, information on Nutrient Reference Values (NRVs) set by FAO/WHO, the European Food Safety Authority (EFSA) and the Institute of Medicine (IOM) are provided on the FAO/WHO GIFT platform to facilitate comparison of the estimated usual intakes of selected nutrients with NRVs. For more details of the statistical program, see (SPADE) and the instruction manual (SPADE - 4100). For NRVs and their applications, see Chapters 8a and 8b, respectively.
The Global Dietary Database (GDD) aims to identify, compile, and standardize individual-level data on dietary factors related to maternal-child health and chronic diseases. Currently the GDD comprises surveys across 188 countries conducted between 1980 and 2018 and provides empirical evidence on dietary intakes both across and within countries worldwide. Priority is given to nationally or sub-nationally representative dietary surveys based on 24-h recalls, food frequency questionnaires or short standardized questionnaires (e.g., Demographic Health Surveys (DHS). Household-level surveys are included if individual-level surveys are not available in a country and converted to individual-level intakes within each household using Adult Male Equivalents (AME), also known as Adult Consumption Equivalents. See: Weisell & Dop (2012). These authors account for the household composition and differing energy intakes by age and sex of household members. See Coates et al. (2017) for discussion of the validity of the application of the AME method for household-level data.
The FoodEx2 categorization system is used to standardize the description and classification of foods into food groups. Standardization also includes categorising nutrients and their units; quality assessment; aggregation by demographic strata and energy adjustment. For more details on data extraction and standardization, see: (GDD)
Mean intakes of 54 dietary factors by country, year, age, sex, education, urbanicity, and pregnancy / lactation status within nations can now be estimated using the GDD prediction model in 188 countries / territories in Asia, Asia-Pacific high-income countries, Oceania, Former Soviet Union, Latin America and Caribbean, Middle East and North Africa, South Asia, Sub-Saharan Africa, and the Western high-income countries in Australasia, Europe and North America (Miller et al., 2021). The dietary factors selected and defined based on evidence for relationships with maternal-child health or chronic diseases include 14 foods, 7 beverages, 12 macronutrients, and 18 micronutrients. See (GDD) for more details on data extraction and the standardization used for estimating dietary intakes. Several indicators of global dietary patterns (e.g., Alternative Healthy Eating Index (AHEI); Dietary Approaches to Stop Hypertension (DASH), and the Mediterranean Diet Score (MED)) among children and adults have also been compiled from the GDD sets and compared globally, regionally, and nationally (Miller et al. ,2022).
Institute for Health Metrics and Evaluation (IHME) initiative database aims to provide rigorous and comparable measurement of the world’s most important health problems and evaluates the strategies used to prevent them. IHME uses FAO Food Balance Sheet estimates, national product sales, household surveys, and data based on 24h recalls (considered the gold standard). Datasets created by IHME are stored in the IHME data catalogue known as the Global Health Data Exchange and can be freely downloaded from the IHME website.
IHME provides freely available modeled data by country, age, sex and year, based on primary data collected in 204 countries and 87 indicators, including 15 dietary indicators (9 foods and 6 nutrients). The dietary indicators are included in GBD 2017, a worldwide observational epidemiological study that tracks the progress within and between countries of the changing health challenges (Lim et al., 2013).


