Pages

Showing posts with label business statistics. Show all posts
Showing posts with label business statistics. Show all posts

Monday, 10 October 2011


Sample and Population

The population consists of the set of all measurements in which the investigator is interested. The population is also called the universe.

A sample is a subset of measurements selected from the population. Sampling from the population is often done randomly, such that every possible sample of n elements will have an equal chance of being selected. A sample selected in this way is called a simple random sample, or just a random sample. A random sample allows chance to determine its elements.

Source: http://mips.stanford.edu/courses/stats_data_analsys/lesson_1/pop_sam.mov

For example, a manufacturer produces 1000 units of a product ‘A’ in a production cycle, then the total number of units in a cycle constitutes a population. Now draw 30 units of the product ‘A’ for a routine quality check; which is called a sample (a miniature or representation of population). If the selection process is random, then the sample can be considered as random sample.

The definition of Population and sample are relative to what we want to consider. If we are dealing with all production cycles in a quarter, then that is our population and total number of production cycles in any two weeks may be considered as sample.

Source: http://mips.stanford.edu/courses/stats_data_analsys/lesson_1/pop.gif

A set of observations/ measurements obtained on some variable is called a data set. For example the units the number of units of product ‘A’ sold through 10 different outlets makes a data set.

A conclusion drawn about a population based on the information in a sample from the population is called a statistical inference. Statistical inference may be based on data collected in surveys or experiments. To ensure the accuracy of statistical inference, data must be drawn randomly from the population of interest, and we must make sure that every segment of the population is adequately and proportionally represented in the sample.

There are challenges in data collection. One of them is non-response bias. This is the biasing of the results that occurs when we disregard the fact that some people will simply not respond to the survey. The bias distorts the findings, because the people who do not respond may belong more to one segment of the population than to another.

In experiments, as in surveys, it is important to randomize if inferences are indeed to be drawn. People should be randomly chosen as subjects for the experiment if an inference is to be drawn to the entire population.

In other situations Data may come from secondary sources like published government statistical abstracts or from any other relevant data providers. In that case, we must get through the background of such data to make use in any realistic analysis. For example, the unemployment rate over a given period is not a random sample of any future unemployment rates, and making statistical inferences in such cases may be complex and difficult.

Usual inspection of the data or information will serve when interest centres on the particular observations. If, however, you want to draw meaningful conclusions with implications extending beyond your limited data, statistical inference is the way to do it.

In marketing research, we are often interested in the relationship between advertising and sales. A data set of randomly chosen sales and advertising figures for a given firm may be of some interest in itself, but the information in it is much more useful if it leads to implications about the background process. The relationship between the firm’s level of advertising, the resulting level of sales, conversion thread in an add, influence variations, impact on brand equity, and split up media planning etc. are some advantage due to ‘Add- Sales- Brand Analysis. An understanding of the true relationship between advertising and sales—the relationship in the population of advertising and sales possibilities for the firm—would allow us to predict sales for any level of advertising and thus to set advertising at a level that maximizes profits.

A quality control engineer at a plant making a product ‘A’ needs to make sure that no more than 2% of the product ‘A’ produced are defective. The engineer may routinely collect random samples of ‘A’ and check their quality. Based on the random samples, the engineer may then draw a conclusion about the proportion of defective items in the entire population of Production ‘A’. These are just a few examples illustrating the use of statistical inference in business situations. 

Friday, 30 September 2011


It is better to be roughly right than precisely wrong.
—John Maynard Keynes

The information from a good statistical analysis is always concise and often precise. Statistics is the science which helps us to summarise, analyse and make inference from data to support fact based decisions in the field of Business, Economics, Healthcare, Education etc. The result of such analytical initiatives will be incredible growth and performance in respective fields. The information (Data) may quantitative or qualitative.

Qualitative data describe items in terms of some quality or categorization.

Eg; If we ask someone’s weight, the answers may be overweight, underweight or obese. The answers are categorical. If we made any query regarding the support from a retail outlet, the result will be good, bad, no so good, and average etc. These types of responses are the marks for quality.

Quantitative data is a numerical measurement expressed not by means of a natural language description, but in terms of numbers for which arithmetic operations such as averaging make sense.

Eg; As it is mentioned in the previous example some can answer his weight as 50kg, or 65kg.  

Statistical inference will deliver insights for strategic decision making. We draws insights from gathered data.  Data generates due to measurement and the measurement should be precise to get dependable results. Accuracy in data collection, careful data entry and cleaning and the use of most suitable analytical tool with clever interpretation will result into best business decisions. So measurement mechanisms should be calibrated continuously and rigorously to get a competitive edge on your decisions.

The four generally used scales of measurement are listed here from weakest to strongest. They are nominal scale, ordinal scale, interval scale and ratio scale.

“Nominal” stands for name of a “category”. The nominal scale of measurement is used for qualitative rather than quantitative data. Here the numbers are used simply for groups or classes.
Eg; Male: Female: 1:2, Good: Bad: Average: : 1: 2: 3

In ordinal scale of measurement, data elements may be “ordered” according to their relative size or quality.
Eg; A product can be ranked by a customer in 5 point scale such that 1 for worst and 5 for best and 2, 3, 4 stands for the opinions between best and worst. Here we do not know how much better one product is than others, only that it is better.


In the interval scale of measurement the value of zero is assigned arbitrarily and therefore we cannot take ratios of two measurements. But we can take ratios of intervals. 
A good example is how we measure time of day, which is in an interval scale. We cannot say 10:00 A.M. is twice as long as 5:00 A.M. But we can say that the interval between 0:00 A.M. (midnight) and 10:00 A.M., which is duration of 10 hours, is twice as long as the interval between 0:00 A.M. and 5:00 A.M., which is duration of 5 hours. This is because 0:00 A.M. does not mean absence of any time.

If two measurements are in ratio scale, then we can take ratios of those measurements. The zero in this scale is an absolute zero. Money, for example, is measured in a ratio scale. A sum of `100 is twice as large as ` 50. A sum of `0 means absence of any money and is thus an absolute zero. We have already seen that measurement of duration (but not time of day) is in a ratio scale. In general, the interval between two interval scale measurements will be in ratio scale. Other examples of the ratio scale are measurements of weight, volume, area, or length. 

will continue...

It's a humble beginning.