Sampling and quantitative analysis

A survey estimate concerns a defined population. Its credibility depends on a suitable sampling frame, documented selection probabilities, consistent measurement and calculations that represent the design. Plan these together, beginning with the groups and decisions the survey must support.

The baseline assessment method sets out the role of survey evidence in appraisal and mitigation. This guide develops sample design and estimation; the questionnaire and fieldwork guide addresses measurement. Apply these choices to the populations, national instruments and field conditions discussed for Maluku and Nepal.

The United Nations household-sampling manual supplies the statistical framework (United Nations, 2008; references). Calculations below illustrate stated assumptions. Complex multistage designs require a survey statistician to establish selection probabilities, variance estimation and attainable precision.

Decide what requires enumeration

An affected-person census identifies potentially displaced persons and their relevant circumstances. The main guide explains why statistical sampling cannot replace that function. A more detailed socioeconomic module may use a designed sample where permitted and suitable, but every household needing an individual assistance decision needs enough information for that decision.

Define three populations separately: the area population to be described, the affected population requiring identification, and any comparison population used for evaluation. Establish the coverage and selection process for each. A comparison population needs a justified relationship to the outcomes the exposed population would have experienced in the project's absence.

Construct the frame before selecting respondents

Create a dated list or mapped area frame. Identify dwellings and their households through field listing; distinguish vacant structures, collective accommodation, businesses, and seasonal use. Include access procedures for remote places and arrangements for people absent during listing. Record additions and removals through an approved frame update.

For stratified household sampling, assign every eligible household to one mutually exclusive stratum before selection. Strata might combine exposure and livelihood conditions where sample sizes permit. They must be exhaustive for the intended population; overlapping categories require a separate classification method. Oversample small groups where their estimates are decision-relevant, and retain inclusion probabilities so the oversampling can be corrected in estimation.

With simple random sampling without replacement within stratum , the inclusion probability is . The base weight is , where is the frame population and the number selected. Equal sample sizes from unequal populations do not produce equal population weights.

Cluster sampling can reduce travel; correlation among nearby households influences precision. For multistage selection, record the probability at each stage. For probability proportional to size, retain the size measure and selection algorithm needed to reconstruct those probabilities. Keep a selection ledger that a second analyst can reproduce.

Plan precision for the decisions

For an estimated proportion under simple random sampling, a conventional planning calculation is

where is the anticipated proportion, is the absolute margin of error, and is the confidence multiplier. With , , and , . This approximation concerns one proportion under the stated design; assess other outcomes and reporting domains separately.

For a finite population and an assumed design effect , one approximate calculation is

If and , this gives approximately 389.52 completed interviews. Round to 390. At an anticipated response rate of 90%, select at least units. The design effect and response rate are illustrative assumptions. In a clustered design, specify finite-population corrections at the relevant sampling stages rather than applying this approximation to the entire frame.

Precision must be evaluated for each reporting domain. A total sample of 390 says little about an outcome among 25 households in a particular livelihood group. Income precision requires an anticipated variance; measuring change requires information about the correlation between rounds; an evaluation requires power calculations for a minimum relevant effect and the intended allocation. Budget more than interviews: listing, translation, revisits, supervision, secure data management, and analysis determine achievable quality.

Treat nonresponse as evidence

Compute response rates from the selection ledger, with the denominator and handling of unknown eligibility explicit. Distinguish unit nonresponse from missing individual items. Examine whether response differs by exposure, location, tenure, language, and other frame characteristics.

Within a suitable adjustment cell, a base-weight response adjustment can be written

where is the set selected and eligible, and is the responding subset. Treat ineligible frame units separately. A cell with no respondents requires further fieldwork or a justified alternative adjustment. The calculation assumes respondents represent nonrespondents conditional on the cell's characteristics; assess that assumption using available frame and response evidence.

Calibration can align weighted totals with reliable population controls. Its interpretation depends on frame coverage, compatible controls and the selection process. Retain base weights, adjustments and final weights separately and report unusually influential weights.

Calculate the right estimand

The weighted mean of a household measure is

For a binary measure, this is the estimated share of households meeting the definition. Define the eligible denominator before calculation. A livelihood indicator among fishing households has a different universe from the same indicator among all households.

Mean household per-capita consumption and consumption per person across the population also differ:

where is household consumption and its eligible membership count. The first gives each represented household equal influence; the second weights household per-capita values by the number of people represented. Label the result accordingly. Roster definitions must match consumption coverage.

For stratified simple random sampling with complete response and fixed frame sizes, the estimated variance of the population mean is

Here is the within-stratum sample variance. The lab implements this design specifically. A stratum with one sampled unit does not support estimation of its within-stratum variance unless it is a certainty stratum with one population unit. This formula should not be reused for cluster samples, arbitrary final weights, or incomplete-response estimates.

Use software that represents the actual strata, sampling stages, weights and finite-population information. R's survey package supports design-based estimation and domain analysis (references). Specify the design for regression and resampling as well as descriptive estimates.

Report distribution and missingness

Present household counts, eligible counts, observed counts, weighted estimates and uncertainty together. State the evidential limits for small subgroups and choose rounding appropriate to the estimate's precision.

For skewed welfare measures, report quantiles and distributions as well as means. Use a survey-aware method for quantile uncertainty. Check zeros, extreme values, unit errors, and duplicated records before deciding how to handle them. If an outlier is valid, show how conclusions change under justified alternative specifications rather than simply removing it.

Missingness can itself be patterned by risk. Tabulate item missingness within exposure groups and by respondent type. Define when a household total is complete; summing five observed items and ignoring three missing items will understate spending. Report complete-case analysis as such. Imputation requires a model, assumptions, and uncertainty treatment; replacing all missing values with a mean narrows variation artificially.

Compare welfare across time

Bring monetary measures to a common price basis:

Use a documented price index with relevant geographic and temporal coverage. A national CPI may incompletely reflect local housing costs or a community's consumption basket. Consider a justified local-price sensitivity analysis. Keep real and nominal values and record index revisions. Do not treat compensation payments or asset sales as recurring livelihood improvement.

For panel change, retain household links and analyze the same measure over a comparable season. Report attrition and compare baseline conditions of retained and lost households. Describe complete-case results as change among the reinterviewed households. Population estimates require an explicit treatment of selective attrition and its assumptions.

Difference-in-differences compares change among an exposed group with change among a comparison group. Its interpretation depends on assumptions including parallel trends in the absence of treatment, stable measurement and appropriate treatment of spillovers. Assess those assumptions with prior observations, contextual evidence and sensitivity analysis. A single baseline supplies no pre-intervention trend (Gertler et al., 2016).

Translate estimates into management

Keep descriptive findings and management judgments distinct. A confidence interval describes sampling uncertainty under assumptions. It does not decide whether exclusion, lost access, or livelihood deterioration is acceptable. A credible individual case can require corrective action even when a group mean changes little.

For each finding, identify the affected group, the evidence, the proposed response, the responsible party, and the date and measure for checking implementation. Use the survey to investigate and improve mitigation, with claims proportionate to the design.