r/AskStatistics 17h ago

Linear regression: many x data points or less points but with replicates?

When building a calibration curve for a process that is linear over most of the observed range, how to determine the best choice between the following options?

(A) increasing the number of values tested, to get more x axis points;

(B) increasing the number of replicates of each measurement –less x points, but more precise estimate of y for each

For example, if an experimental setup lets me run 12 measurements for a linear calibration curve, is it better to run 4 values in triplicate? Or 6 in duplicates?

4 Upvotes

5 comments sorted by

2

u/Appropriate-Yak001 16h ago

If you believe the process is linear over the range of interest, then using 4 values in triplicate may make more sense than 6 values in duplicate, because the additional replicates give you a better estimate of measurement variability at each x-value. However, the best design also depends on the actual spacing of the x-values and the standard deviation of the measurements.

1

u/babebiboba 13h ago

That makes sense, do you have a thought on how to determine the best design from first principles? i. e. What would be the process to make an educated choice if I have access to an arbitrary number of measurements? My intuition is indeed that there is a qualitative difference between 3 and 4 (4 allows you to identify which point is an outlier, which you can't do with 3) but that that it's always better to stop at 4 and max the accuracy of those 4 points, but I might be wrong

2

u/Temporary_Stranger39 16h ago

How certain are you of the shape of the curve?

1

u/babebiboba 13h ago

Let's assume pretty certain, things like "light absorbance is proportional to concentration of molecule" – linear for all intents and purposes

1

u/Educational-Paper-75 2h ago

If the variance is assumed constant you really don't need many replicates.