Pathors
Statistics GuideOct 6, 2026

Statistics You Need Before Running AI Phone Surveys: 1,068 Completes, Confidence Levels, Margin of Error and Interviewer Effects (2026)

Pathors Team

Pathors Team

Pathors

Statistics You Need Before Running AI Phone Surveys: 1,068 Completes, Confidence Levels, Margin of Error and Interviewer Effects (2026)

"1,068 completed interviews, ±3% margin of error at the 95% confidence level." The line shows up on customer satisfaction studies, market research and other telephone survey reports, yet few readers really understand it. When the interviewer changes from a person to an AI, these concepts do not change; they matter more. AI phone surveys change where error comes from and what shape it takes, not the rules of statistics.

This is a tutorial. We walk through seven concepts you need before running an AI phone survey or designing a telephone questionnaire, each with a worked business-survey example, so you can judge how far a phone survey's numbers can be trusted.

Concept 1: Confidence level and margin of error, and where 1,068 comes from

To estimate a proportion (say, the share of satisfied customers) under simple random sampling, the required sample size is:

n = z² × p(1−p) ÷ e²

  • z reflects the confidence level: 1.96 for 95%, 2.576 for 99%
  • p is the expected proportion; use 0.5 when unknown, because p(1−p) peaks there and gives the most conservative answer
  • e is the acceptable margin of error
  • Worked example: 95% confidence, p = 0.5, e = 0.03 gives n = 1.96² × 0.25 ÷ 0.0009 = 1,067.1, rounded up to 1,068.

    "95% confidence" means that if you repeated the sampling many times with the same method, about 95% of the intervals would contain the true value. It does not mean the true value has a 95% probability of lying in this particular interval. Keeping ±3% at 99% confidence needs about 1,844 completes.

    Also, ±3% is the maximum, at p = 0.5. If a result is 10%, the same 1,068 completes give about ±1.8%.

    Concept 2: The sample-size table and subgroup error

    Sample sizeMax margin of error (95% confidence)
    100±9.8%
    200±6.9%
    300±5.7%
    500±4.4%
    1,068±3.0%
    1,500±2.5%
    2,401±2.0%

    Two takeaways: error shrinks with the square root of the sample, so halving it takes four times the sample; and ±3% belongs to the full sample only.

    Worked example: in a 1,068-complete customer satisfaction survey, only 200 respondents became customers in the past year. Their subgroup error is about ±6.9%. If the report says "new customers 52% satisfied vs 48% overall", that 4-point gap sits inside the subgroup error; you cannot conclude new customers are happier.

    Concept 3: Response rates, and why you dial many more numbers

    1,068 is the number of completes, not dials. Common funnel measures are the contact rate (working numbers where a respondent was reached), the cooperation rate (contacted eligible respondents who completed) and the response rate (completes relative to all eligible cases). AAPOR's Standard Definitions are the widely used reference.

    Worked example (illustrative; rates are assumptions, not industry data):

    StageAssumed rateNumbers
    Numbers dialled—12,715
    Working numbers70%~8,900
    Eligible respondent reached40%~3,560
    Interview completed30%~1,068

    Under these assumptions you dial about 12 numbers per complete, before callbacks. A low response rate does not by itself make results wrong, but it raises the next question: are the people who did not answer like the people who did?

    Concept 4: Sampling error vs non-sampling error

    Sampling error is random error from interviewing only part of the population; a formula captures it and it shrinks as the sample grows. Non-sampling error is systematic and does not go away with a bigger sample:

    TypeIn plain wordsExample
    Coverage errorSome people are not in the frameDialling landlines only misses mobile-only customers
    Nonresponse errorResponders differ from refusersBusy working customers complete less often
    Measurement errorAnswers do not reflect true viewsInterviewer tone, wording, mishearing, recording mistakes
    Processing errorMistakes in handling dataManual entry and coding errors

    The ±3% on a report covers only the first kind of error, sampling error. When two surveys on the same topic differ by 8 points, the reason is usually in this table.

    Concept 5: Interviewer effects, and what changes when the interviewer is an AI

    When different interviewers administer the same questionnaire, answers vary systematically by interviewer. This interviewer effect can be measured with an intra-interviewer correlation ρ, and it inflates variance:

    Design effect ≈ 1 + ρ × (average completes per interviewer − 1)

    Worked example: 20 interviewers share 1,068 completes (about 53 each). Assume ρ = 0.01 (illustrative). The design effect is about 1 + 0.01 × 52 = 1.52, the effective sample drops to about 701, and the margin of error widens from ±3.0% to about ±3.7%, while the report still says ±3%.

    With an AI interviewer, every call uses the same wording, tone and pace, read verbatim, so the between-interviewer component disappears. That is a statistical property, not a marketing claim: consistent wording means consistent measurement conditions.

    The flip side: every call shares one "interviewer". If the script leads, the bias does not average out; it applies identically to every call. AI turns interviewer variance into a question of whether the script is neutral, which makes piloting and question review more important. Speech-recognition errors and mode effects (people answering a machine differently) are measurement error too, and need validation.

    Concept 6: Nonresponse bias and weighting: what weighting fixes and what it can't

    Nonresponse bias is roughly the nonresponse rate × the difference between responders and non-responders.

    Worked example: response rate 30%; 60% of responders are satisfied, but only 50% of non-responders are. The true value is 0.3 × 60% + 0.7 × 50% = 53%, yet the report says 60%: a 7-point bias, far larger than ±3%.

    Weighting is the usual correction: align the sample's structure with the population. Worked example: customers aged 60+ are 40% of the customer base but 30% of the sample, so their weight is 40 ÷ 30 ≈ 1.33; everyone else gets 60 ÷ 70 ≈ 0.86.

    Weighting can fix imbalances on the weighting variables (age, sex, region). It cannot fix differences between willing and unwilling respondents within the same group, or people the frame never covered. It also has a cost: the more uneven the weights, the smaller the effective sample. The weights above reduce 1,068 completes to an effective ~1,019; more variables and more extreme weights inflate error further.

    Concept 7: Question wording and order effects in voice surveys

    People behave differently when they hear a questionnaire instead of reading it:

  • Recency effect: on the phone, respondents cannot see the options and tend to pick the last one they heard. Keep options few (five or fewer) and rotate their order across respondents.
  • Order and context effects: earlier questions colour later ones. Asking "Did you wait long the last time you called support?" before "Overall satisfaction" can pull satisfaction down.
  • Acquiescence: "Do you agree that…?" invites "yes". Use balanced wording: "Do you lean more towards A or B?"
  • Read the full scale: reading only the end points pushes answers to the extremes or the middle.
  • These effects exist with human and AI interviewers alike, but AI repeats your design word for word: a good design is good on every call, a biased one is biased on every call.

    Compliance note (Taiwan)

  • Opinions and demographics collected in a phone survey are personal data. Under Article 8 of the PDPA, tell respondents who is collecting, the purpose, the data categories and how the data will be used; non-government collectors need a specific purpose and lawful basis (Article 19) and appropriate security measures (Article 27).
  • We recommend also stating up front that the interviewer is an AI, and setting retention periods for recordings and raw data. For other outbound compliance points, see Is AI Outbound Calling Legal in Taiwan?
  • Put the seven concepts together: ±3% describes sampling error only. Whether a phone survey can be trusted depends on coverage, nonresponse and measurement error, and on what weighting costs. An AI interviewer has one clear statistical advantage on measurement: it applies the same script verbatim on every call, removing differences between interviewers, and it keeps a recording and transcript of each call so what was asked and answered can be audited. Sampling, weighting and questionnaire design remain the researcher's job.

    Further reading:

  • Optimizing AI Voice Agents: A/B Testing and Script Iteration
  • Is AI Outbound Calling Legal in Taiwan? PDPA, Recording & Debt-Collection Rules
  • Evaluating AI outbound tools? Start with How to Choose an AI Auto-Outbound Calling System.
  • Frequently Asked Questions

    Why do phone surveys so often use 1,068 completed interviews?

    At 95% confidence, a conservative expected proportion of 0.5 and a ±3% margin of error, n = 1.96² × 0.5 × 0.5 ÷ 0.03² = 1,067.1, which rounds up to 1,068.

    What does a 95% confidence level mean?

    If you repeated the sampling many times with the same method, about 95% of the confidence intervals would contain the true value. It does not mean this particular interval has a 95% chance of containing it.

    Does a ±3% margin of error mean the result is accurate to within 3 points?

    No. ±3% covers sampling error only, and only for the full sample. Coverage error, nonresponse bias, interviewer effects and question wording are not included, and subgroups have larger error.

    Are AI phone surveys more accurate than human interviewers?

    AI does not change sampling error, but reading the same script verbatim on every call removes differences between interviewers. The trade-off: a biased script biases every call, and speech recognition and mode effects need validation. Compare against your existing method.

    Can weighting correct every bias?

    No. Weighting only corrects imbalances on the weighting variables such as age, sex and region. It cannot fix differences between responders and refusers within a group, and it reduces the effective sample size.

    How many completes does a customer satisfaction phone survey need?

    It depends on the precision you need and how many subgroups you will compare. At 95% confidence, ±5% needs about 385 completes and ±3% about 1,068. If you compare branches or customer segments, each subgroup also needs enough completes of its own.

    Pathors Team

    Pathors Team

    Pathors

    Passionate about leveraging AI technology to transform customer service and business operations.

    Read More Articles

    Automate Every Conversation That Matters.