Showing posts with label MR9. Show all posts
Showing posts with label MR9. Show all posts

Sunday, May 18, 2008

Problems with sample data (MR, unit 9)

Potential problems:
  • Bias - some households may have poor chance of being selected in sampling frame is out of date, individuals decline to respond, bias in questionnaire or interview
  • Insufficient data - sample too small to be reflect whole population
  • Unrepresentative data - data might be collected in abnormal conditions
  • Omission of important factor - important questions omitted in design of questions
  • Carelessness
  • Confusion of cause & effect - wary of assuming associated variables mean one causes the other
  • Interpretation - true of depth interviews with lengthy replies
Solving problems (ANN):
  • Accuracy - insert control questions into questionnaire. Reply to one should be compatible with another. If not, the value of responses could be dubious or interviewee may be confused.
  • Non-sampling error (results from way observations made) - poorly worded questions, unsuitable people surveyed, purpose of study may affect responses, personal bias of interviewer, non-response
  • Non-response - units outside population (demolished houses), unsuitable for interview, movers, refusals, away from home
  • common in random sample surveys
  • rising crime and data fatigue reduces response rates (sugging & frugging - selling/fundraising under guise of research)
  • can be avoided (except in mail surveys):
    • DO NOT JUST APPROACH NEIGHBOUR, affects sampling framework
    • People who've moved. Select individual from new household with rigorous procedure.
    • Minimise refusals by keeping brief as possible using skilled interviewers and financial incentives
    • Contact people away from home at later date. Plan call times sensibly as most out at work during day.

Saturday, May 17, 2008

Sample sizes using statistics (MR, unit 9)

Can only be used for probability samples.

Need 3 pieces of information:
  • O = estimate of standard deviation of population
  • Z = level of confidence expressed in standard errors
  • R = acceptable level of precision (the sampling error)
N = sample size
SE = standard error of the mean

MEASURING AVERAGES


N = (Z * O /R)2
sample size = (2.58*£50/£10)2

This is just using algebra to calculate n in the equation to find the standard error.

Example: We estimate average salary to be £20,000 with a standard deviation of £50. What size of sample should be taken to estimate the true average salary to within £10, with 99% confidence?£10 = 2.58*£50/ route sample size
n = (2.58*50/10)2
n = 166.41 (round up to 167)
A samples size of 167 needed to estimate the true average income to within £10 and be 99% confident of our answer.

MEASURING PROPORTIONS

N = Z2 * PQ / R2
How many people should be sampled to be 99% confident of the number of people likely to take up offer within 0.5% of the population?
P = 20% 0.2
R = 0.5% 0.005

n = 2.58(2) * 0.2 * (1-0.2) / 0.005(2)
n = 5,219

Statistics and sampling (MR, unit 9)


Sampling distribution of the mean
Calculate mean of each sample.
Count number of times each value occurs and plot results as a distribution.

Properties:
  • Results are very close to the normal distribution, a statistical rule called the central limit theorem. Larger samples are obviously even closer.
  • The mean of the sampling distribution (u) = the population mean
  • The sampling distribution has a standard deviation called the standard error (the dispersion of values around the mean).

Standard error of the mean:
s is the sample standard deviation (the sample based estimate of the standard deviation of the population, as this is usually not known)
n is the size (number of items) of the sample.
Confidence levels, limits & intervals
68% of population lies with sample mean +/- 1 SE
95%
of population lies with sample mean +/- 1.96 SE
99%
of population lies with sample mean +/- 2.58 SE

68%, 95%, 99% = confidence levels (degrees of certainty)
Edge of ranges =
confidence limits
The ranges themselves =
confidence intervals

Sampling distribution of a proportion

Many surveys deal not with a population means but with proportions (attitudes, or the percentage of time an event occurs).
Arrange as you would for sampling distribution of a population. Then use following calculation

Standard error of a proportion =

p = the proportion in the sample (taken as an estimate of the population)
q = 1-p
n = size of sample

Example: 285 people out of sample of 400 regularly travel by bus
se = square route (pq/n)
se = square route (285/400 * 1-0.7125 / 400)
se = 0.0226
99% confidence interval = 2.58*se
0.7125 +/-(2.58*0.0226)
With 99% confidence we can say that between 65-77% of people regularly travel by bus

Sample sizes (MR, unit 9)

No universal law for determining, rely on experience
Larger the sample size, the more accurate the results
Point at which increasing sample size has little impact on results

Consider:
  • money & time available
  • degree of precision required (e.g. road widening scheme v. drug testing)
  • number of subsamples required - stratified sampling needs large overall size to ensure adequate representation of each strata
Minimum sample sizes: 300-500
Less than 50 ONLY for qualitative data
Any subgroup of less than 100 may not be statistically valid
National surveys for consumer goods use sample size: 1500-200

Sampling (MR, unit 9)

Population: set of individuals or items from which a statistical sample is taken
Sampling: one of most important marketing research tools because population often too large to take complete survey (census)
Census: survey entire population
Once certain sample size reached, very little accuracy in examining more
Higher cost of census may exceed value of results
Changeability - census data often out of date by time it's collected

Choosing a sample
Must be complete (cover all relevant aspects of population to be examined) or will be biased.
3 types of sample:
  1. Random sampling
  2. Quasi-random sampling: systematic, stratified, multistage
  3. Non-random sampling: convenience, quota, cluster
1) Random sampling
A simple random sample is selected in way that every item in the population has an equal change of being included
Not necessarily a perfect sample - only census would eliminate all chance of bias

Sampling frame required in random sampling: a numbered list of all the items in the population
Should be:
  • complete - include all members of population
  • accurate
  • up-to-date
  • convenient - accessible
  • without duplication
Problems: can be too expensive

2) Quasi-random sampling: good approximation to random sampling
  • Systematic sampling
Selects every nth item after a random start
e.g. sample of 20 from population of 800 (800/20 = 40). 40 = sampling interval. Choose every 40th item after random start.
Must ensure no regular pattern to the population which if it coincided with sampling interval would lead to bias. Avoided with multiple starting points and sampling intervals.

  • Stratified sampling (often best method)
Population must be divided into strata (categories)
Removes possibility of sample being all from same demographic, or location. Random samples taken from each strata - in proportion to weight of each strata in the population.
Problem: requires prior knowledge of each item in the population.

  • Multistage sampling
Divide population into large groups (usually on geographic basis). Small sample of groups taken at random. These are subdivided into smaller groups. A number is selected at random. Process repeated as many times as necessary. Random sample of individuals then taken in each of the smallest groups.
Problem: an approximation to a random sample


3) Non-random sampling: used when a sampling frame cannot be established
  • Convenience sampling
Internet & e-mail based surveys use. Ask the most accessible members of the population. Useful for exploratory if composition of selected sample is reasonably similar to the population of interest. Cheap & simple. Problem: Not precise. Unlikely to reflect population as whole. NB. there's no reliable sample frame of email addresses. Becoming increasingly prevalent but unresolved concerns over its representativeness.
  • Quota sampling
Like convenience but interviewers interview everyone they meet up to certain quota. Partly overcome bias by subdividing the quota into different types of people to ensure the sample measures the structure of the population.
Problem: prone to bias

  • Cluster sampling
Clusters of population units are selected at random (each cluster intended to represent the total population in microcosm) e.g. random phonebook selections. Then some of these clusters are studied. Only need to produces sampling frame for those clusters.
Unlike stratified sampling where subsets have to be homogenous and significantly different from each other. Cluster can therefore be used if adequate sampling frame not available.
Problem: limited to situations where pop can be easily divided into representative clusters.