DOI : 10.5281/zenodo.23256508
- Open Access

- Authors : Kritisha Sachan, Dr. Saravanakumar R
- Paper ID : IJERTV15IS100249
- Volume & Issue : Volume 15, Issue 10 , October – 2026
- Published (First Online): 09-10-2026
- ISSN (Online) : 2278-0181
- Publisher Name : IJERT
- License:
This work is licensed under a Creative Commons Attribution 4.0 International License
Separating Seasonal Bias from Long-Term Drift in Low-Cost PM2.5 Sensors: A LabVIEW Virtual Instrument and Season-Matched Evaluation Using Collocated Data from Bengaluru and Delhi
Kritisha Sachan
Department of Electronics and Instrumentation Engineering, Vellore Institute of Technology (VIT), Vellore, Tamil Nadu, India
Dr. Saravanakumar R
Department of Control and Automation Vellore Institute of Technology (VIT), Vellore, Tamil Nadu, India
Abstract – A study about cost optical sensors that measure tiny particles in the air is being shared. These sensors are becoming more common in India to help with monitoring. Its not clear how well they work over a long time in hot and humid places. This is because changes in humidity and the type of particles in the air can make the sensors look like they are getting worse even if they are not. This paper talks about a computer program made with LabVIEW. This program takes data from sensors and accurate measuring tools that are next to each other. It calculates how much the sensors are off how much they change each month and other details. The program also shows trends in the errors and a way to check for changes over time. Based on data from one source, the study investigated two specific cities: Bengaluru and Delhi. In Bengaluru the difference between the sensor and the real measurements changed a lot during the year. When the data was compared for the same time of year the changes were small. In Delhi the difference between the sensor and the real measurements got bigger over time. This happened in every level of pollution. A calibration done in the year didn’t work as well in the second year. The results from the LabVIEW program match with another program made with Python. The study suggests that sensors need to be checked and adjusted every year in cities, with a lot of pollution.
Keywordslow-cost sensors, PM2.5, PurpleAir, sensor drift, virtual instrumentation, LabVIEW, calibration, India
-
INTRODUCTION
Fine particulate matter (PM2.5) is one of the environmental risks to human health. In India the situation is especially concerning because ambient
PM2.5 levels are very high while the network for monitoring air quality is limited. Cost optical sensors offer a practical way to fill in the gaps in spatial coverage. The PurpleAir PA-II, which uses two laser-scattering Plantower sensors is one of the widely used devices in this category. But theres a problemits data cannot be directly compared to measurements from reference instruments. Studies in the United States have shown that raw PurpleAir readings tend to overestimate PM2.5 by around 40 percent in areas. Applying a correction factor that accounts for humidity helps reduce this error significantly. Other field studies, in environments have found that sensor readings change with the seasons and that individual units can drift over time. Ongoing research continues to compare low-cost sensors and assess different correction methods with new findings being published regularly.
In India, Campmier and colleagues [1] placed PurpleAir PA-II sensors next to beta-attenuation monitors (BAMs) in Delhi, Hamirpur and Bengaluru. They found that calibrations that are specific to the location and the season make the measurements more accurate. They also reduced the errors that happen during times of the day and different seasons. Their main goal was to focus on calibration.. There is another important question that still needs to be answered. Does the error from the sensor change over time because the sensor gets older or does it only change because the air around it changes? With data that covers one to two years it is hard to tell the difference. That is because the same sensor experiences a monsoon, a winter and a summer. Each of these seasons has humidity and different types of particles, in the air. So when someone fits a line to the monthly error they can’t say for sure that the error is changing because the sensor is aging.
This paper looks at that question using the public data archive that comes with reference [1]. Offers four main contributions. First the paper shows a LabVIEW instrument (VI) that calculates monthly performance metrics and trend indicators, from collocated data and displays them on an interactive front panel. Second the paper provides a seasonmatched drift comparison with dayblock bootstrap confidence intervals, that compares the calendar months in successive years. Third the paper applies this method to two different Indian sites revealing that drift depends on the site. Fourth the paper conducts a calibrationtransfer test that demonstrates the cost of reusing a calibration after one year.
-
DATA AND STUDY DESIGN
-
Dataset
We used the open data archive released with [1] (Dryad, doi:10.6078/D1RQ70) [6]. For each site it gives hourly readings of reference PM. from a BAM, PurpleAir PM. under three processing algorithms (CF1, ATM and ALT), and the PurpleAir’s own relative humidity (RH) and temperature. We worked with the CF1 values because our VI uses that channel, and with the Bengaluru
and Delhi sheets only. The Hamirpur record (3,639 h, March 2020 to January 2021) covers less than a year, so a same-season comparison across years is not possible and we left it out. Readings are stamped at the half hour. Within each sheet the rows are in time order, with no missing or repeated entries, and we did not filter the data any further. The dew-point column was ignored because its values (roughly 40 to 80) cannot be in °C, as the archive states. Table I summarises the records.
-
Performance quantities
For hourly sensor values si and reference values ri, over
a set of n hours, we use
b = (1/n) (si ri), = si / ri (1) RMSE = [ (1/n) (si ri)2 ]1/2, NRMSE = RMSE /
mean(r) (2)
r = m s + c (ordinary least squares, reference regressed on sensor) (3)
where b is the mean bias, the ratio of means (a scale- free measure of over-reading) and m, c the fit slope and intercept. Both b and are reported because the bias depends on the pollution level, whereas does not.
TABLE I. Dataset summary and uncalibrated performance (hourly data, CF1 channel)
Quantity
Bengaluru
Delhi
Period of record
21 Jun 2019 31 Jul 2020
24 Jul 2018 3 Jan 2020
Hourly sensorreference pairs, n
7,441
7,504
Mean BAM PM2.5 (g m3)
25.4
104.2
Mean PurpleAir CF1 PM2.5 (g m3)
38.1
162.1
Ratio = mean(CF1)/mean(BAM)
1.50
1.56
Mean bias (g m3)
+12.7
+58.0
MAE (g m3)
13.7
69.0
RMSE (g m3)
18.9
93.2
NRMSE (RMSE / mean BAM)
74.5 %
89.5 %
R2
0.834
0.784
Fit slope m (BAM on CF1)
0.504
0.526
Fit intercept c (g m3)
6.16
18.96
-
-
METHODOLOGY
-
LabVIEW virtual instrument
The V (Fig. 2) reads a comma-delimited file through Read Delimited Spreadsheet (double-precision output), transposes the array and indexes three columns: month number, sensor and reference. An outer For loop runs once per month (N = 14). Inside it, an inner For loop compares every rows month number with the loop counter and passes the matching sensor and reference values through conditional tunnels, which yields that months data. For each month the VI evaluates Linear Fit
(slope, intercept), the mean of the sensorreference difference (bias) and the square root of the mean squared difference (RMSE). Auto-indexed output tunnels collect these as 14-element arrays, which feed waveform graphs of bias and slope against month and numeric indicators. A second Linear Fit of the bias array against the month index gives the overall bias trend. Finally, Index Array selects months 1, 2, 13 and 14 (JuneJuly of the two years) to compute the same-season drift index.
Dss = [ (b13 + b14)/2 (b1 + b2)/2 ] / 12 (g m3 per month) (4)
A scatter plot of sensor against reference for all hours is drawn on an XY graph (Fig. 3). The VI as implemented processes the Bengaluru record; the Delhi record and
the additional statistics below were computed with Python (NumPy, SciPy, stats models), using the same definitions.
Fig. 2. Block diagram of the virtual instrument: file reading and column indexing (bottom), nested For loops for per- month selection and statistics (centre) and trend and same-season drift calculations (right).
Fig. 3. Front panel of the virtual instrument for the Bengaluru record: monthly bias, RMSE, slope and intercept arrays, overall bias trend and same-season drift indicators, bias and slope graphs, and the sensor-versus-reference scatter plot.
-
Verification of the VI
We checked the 56 monthly values the VI displays (bias, RMSE, slope and intercept for each of the 14 months) against a separate calculation in Python with SciPy. The largest differences were 4.9×10 g m³ for bias,
4.7×10 g m³ for RMSE, 4.8×10 for slope and
1.5×10 g m³ for intercept. These are smaller than the precision shown on the front panel. The overall bias trend (0.350784 g m³ per month) and the same-season drift
index (0.0386884 g m³ per month) agreed with Python to every digit displayed.
-
Trend tests
We applied the MannKendall test [7], [8] and Sen’s slope [9] to each monthly series that had at least 150 valid hours (14 months for Bengaluru, 15 for Delhi). These tests show whether a steady upward or downward trend exists, but they do not remove the effect of the seasons.
-
Season-matched comparison
To separate drift from seasonal change, we compared the same calendar months in consecutive years, using only months with enough data in both years. For Bengaluru this was JuneJuly (2019 vs 2020; 897 and 693 h), and
for Delhi OctoberDecember (2018 vs 2019; 1,699 and 1,382 h). Drift was taken as the year-2 value minus the year-1 value, for both and b. We obtained 95 % confidence intervals with a day-block bootstrap of 2,000 replicates, which resamples whole days so that the correlation between hours within a day is kept. To make sure concentration was not driving the result, we repeated the comparison within bands of reference concentration.
-
Calibration-transfer test
We fitted a linear correction, r = a + b s + c RH, to the matched months of year 1, following the humidity-
dependent corrections in [2]. Its accuracy within that year was estimated by five-fold cross-validation, with folds split by day. We then applied the same model, unchanged, to the same months of year 2. If the error rises from the cross-validated value to the year-2 value, the relationship between sensor and reference has changed.
-
Covariate-adjusted regression (sensitivity check)
As an alternative approach, we regressed ln(s/r) on RH, temperature, ln r and elapsed time t in months, with and without one annual harmonic. Standard errors were corrected for unequal variance and autocorrelation (NeweyWest, 48 h lag) [11].
-
-
RESULTS
-
Overall Performance
Without calibration, the sensor reads high at both sites (table i). The mean cf1 value is 1.50 times the bam value in bengaluru and 1.56 times in delhi. The sensor and reference follow the same pattern closely (r² = 0.83 and 0.78), but the values are far apart (nrmse 75 % and 89 %). The fitted slope of about 0.5 means that two sensor units correspond to roughly one reference unit. The sensor is therefore consistent but biased, in line with earlier findings [1], [2]. Fig. 4 shows that this relationship also changes with the season.
Fig. 4. Hourly PurpleAir CF1 against BAM PM2.5 for (a) Bengaluru and (b) Delhi, coloured by season. Dashed line: 1:1
-
Monthly behaviour
I read Table II and see that VIs monthly outputs for Bengaluru show a pattern. Bias rises steadily from +1.80 g m3 in June 2019 to +23.74 g m3 in December 2019 then drops sharply to 0.01 g m3 by June 2020. RMSE tracks Bias closely starting at 4.61 g m3 and climbing to 28.73 g m3. shifts from 1.0 during the monsoon months to 1.78 in December. Slope values vary between 0.31 and 0.65. In Delhi the seasonal cycle is even stronger as shown in Fig. 5B. Monthly in Delhi
stays around 1.82.0 from September to February falls below 1 in MayJune 2019 and rises again to 1.41.6 in the following autumn. The ratio does not simply follow humidity. In Bengaluru the ratio peaks in November January after RH has fallen from its monsoon maximum. In Delhi, the lowest relative humidity (RH) during MarchMay coincides with a drop in the ratio. The ratio peaks again, in winter when RH is moderate. These patterns show that pollution level and aerosol composition also play a role as illustrated in Fig. 5.
TABLE II. Monthly performance of the PurpleAir CF1 sensor in Bengaluru (VI outputs, 14 months)
Month
Month
n (h)
Bias (g m3)
RMSE (g
Ratio
Slope m
Intercept c
no.
m3)
1
Jun 2019
208
1.80
11.96
1.08
0.315
14.82
2
Jul 2019
689
2.59
7.06
1.15
0.389
9.75
3
Aug 2019
706
2.68
7.46
1.15
0.434
8.67
4
Sep 2019
653
7.41
12.73
1.39
0.464
6.87
5
Oct 2019
654
12.39
17.05
1.60
0.600
0.80
6
Nov 2019
648
22.93
27.64
1.73
0.553
1.32
7
Dec 2019
682
23.74
26.68
1.78
0.520
2.18
8
Jan 2020
680
23.19
28.73
1.62
0.521
5.71
9
Feb 2020
423
20.51
23.51
1.57
0.517
6.64
10
Mar 2020
439
p>14.69 18.50
1.48
0.524
6.88
11
Apr 2020
646
16.26
19.13
1.56
0.540
4.55
12
May 2020
320
11.20
14.14
1.40
0.530
7.08
13
Jun 2020
377
0.01
4.61
1.00
0.651
5.95
14
Jul 2020
316
5.33
9.23
1.35
0.466
5.74
Fig. 5. Monthly mean sensor/reference ratio (solid, left axis) and mean RH (dashed, right axis) for (a) Bengaluru and (b) Delhi. Open circles: months with fewer than 150 valid hours.
-
Effect of humidity
When we split the Bengaluru data by relative humidity the value of rises steadily from 1.29 when relative humidity is less than 40 percent to 1.68 when relative humidity is greater than 70 percent. The NRMSE grows from 40 percent to 96 percent. The fit slope drops from
0.62 to 0.47. This matches the idea that hygroscopic particle growth makes the optical signal larger, which’s why the humidity term appears in reference [2]. In Delhi
relative humidity goes above 70 percent in one hour of the record. The value of increases with relative humidity over the 529-hour period: it stands at 1.34 below 40% RH, rises to 1.71 between 40% and 55% RH,
and reaches 2.06 between 55% and 70% RH. Because relative humidity changes with season and concentration these results show relationships rather, than pure humidity effects.
TABLE III. Bengaluru performance by relative-humidity band
RH band
n (h)
Mean BAM
Ratio
Bias
RMSE
NRMSE
Slope m
<40 %
1165
28.0
1.29
+8.0
11.3
40 %
0.62
4055 %
1330
27.4
1.37
+10.1
15.7
57 %
0.56
5570 %
2202
25.8
1.51
+13.2
19.7
76 %
0.50
>70 %
2744
23.0
1.68
+15.6
22.0
96 %
0.47
Bias and RMSE in g m3; mean BAM PM2.5 in g m3.
-
Trend tests
Neither bias nor rho shows a monotonic trend at either site (Table IV): in Bengaluru the monthly bias trend is
+0.44 microgram per cubic meter per month (95 percent confidence interval from -1.74 to +1.87; p equals 0.59) and in Delhi rho shows a non-significant decline (p
equals 0.093). The fit slope m does rise significantly at both sites (p equals 0.026 and less than 0.001). However m is estimated from within-month regressions and depends on the concentration range, in each month so it is a drift indicator when concentrations are seasonally variable; this is examined next.
TABLE IV. MannKendall and Sens slope tests on monthly series
Site
Monthly series
Kendall
p
Sens slope per month (95 % CI)
Bengaluru
Bias (g m3)
0.12
0.591
+0.440 (1.744 to 1.870)
Ratio
0.03
0.914
+0.007 (0.046 to 0.051)
Fit slope m
0.45
0.026
+0.012 (0.001 to 0.022)
Delhi
Bias (g m3)
-0.03
0.923
0.697 (7.907 to 9.841)
Ratio
-0.33
0.093
0.038 (0.091 to 0.009)
Fit slope m
0.64
<0.001
+0.015 (0.008 to 0.032)
-
Season-matched drift
Delhi looks different. Between OctoberDecember 2018 and the same months in 2019, fell from 1.86 to 1.51 ( = 0.35, 95 % CI 0.40 to 0.29), and the mean bias fell by 39.9 g m³. The drop appeared in every month (October 1.78 1.63, November 1.88 1.54, December 1.91 1.38) and in all five concentration bands from 4080 to 200300 g m³ (Fig. 6c; for example, 1.89 1.68 at 80120 and 1.80 1.49 at 200
300 g m³). Mean RH was slightly higher in 2019 than in 2018 (42.6, 41.8 and 43.7 % against 37.9, 38.6 and 40.5 % for the three months). An increase in humidity would expectedly elevate ; therefore, neither concentration nor humidity explains the observed reduction. Thus over one year the Delhi sensors response relative to the BAM declined by 19 % (from 1.86 to 1.51) while the fit slope increased from 0.53, to 0.71.
TABLE V. Season-matched comparison of the same months in consecutive years
Site
Months
n (h), yr 1 / yr 2
Ratio , yr 1
yr 2
(95 % CI)
Bias, yr 1 yr 2
bias (95 % CI)
Bengaluru
JunJul
897 / 693
1.13 1.15
+0.02 (0.06 to 0.11)
2.4 2.4
+0.0 (1.3 to
1.5)
Delhi
OctDec
1699 / 1382
1.86 1.51
0.35 (0.40 to
0.29)
118.0 78.1
39.9 (51.4 to
28.9)
Fig. 6. Season-matched comparison of the mean sensor/reference ratio: (a) Bengaluru, JuneJuly 2019 vs 2020; (b) Delhi, OctoberDecember 2018 vs 2019; (c) Delhi, OctoberDecember pooled, by BAM concentration band.
-
Calibration transfer
Table VI shows what the change means for a user who calibrates once and then uses the model again. In Delhi the model that was fitted to October to December 2018 has a -validated NRMSE of 15.4 percent (R2 equals 0.90) but when it is used without any changes for October to December 2019 its NRMSE is 33.4 percent (R2 equals 0.72) with a mean under-prediction of 29.2 microgram per cubic meter about 19 percent of the mean BAM value
(153.3 microgram per cubic meter). In Bengaluru NRMSE went up slowly from 27.3 percent to 32.6 percent with a mean over-prediction of 3.4 microgram per cubic meter (about 21 percent of the mean BAM value of 16.3 microgram per cubic meter); at these low concentrations small absolute errors are large, in relative terms and cross-validated R2 is low (0.39).
TABLE VI. Calibration fitted to year-1 matched months and applied unchanged to year 2
Site
Year-1 model (BAM = …)
NRMSE, yr- 1 CV
R2, yr-1 CV
NRMSE, yr 2
R2, yr 2
Mean error, yr 2 (g m3)
Bengaluru
21.25 + 0.426 CF1 0.173 RH
27.3 %
0.39
32.6 %
0.27
+3.4
Delhi
6.03 + 0.533 CF1 0.124 RH
15.4 %
0.90
33.4 %
0.72
29.2
-
Covariate-adjusted regression
In Delhi, the regression etimate of the time coefficient is 0.021 per month in ln(s/r) (95 % CI 0.029 to
0.012), and 0.022 (0.028 to 0.015) with an annual harmonic, in agreement in sign and approximate size with the season-matched result. In Bengaluru, however, the same regression yields positive coefficients, +0.026 (0.016 to 0.036) without and +0.018 (0.011 to 0.025) with the harmonic, which contradicts the season-matched finding of no change. The record there spans only 13.3
months, so elapsed time and the annual cycle are nearly collinear, and the non-sinusoidal seasonal pattern of (Fig. 5a) is absorbed by the time term. We therefore regard the Bengaluru regression coefficient as unidentified rather than as evidence of drift. The stated intervals, which allow only 48 h of autocorrelation, are also likely to be too narrow. This discrepancy illustrates why regression or trend estimates on records shorter than two annual cycles should not be read as drift, and supports the matched-season design.
-
-
DISCUSSION
Two sites, two outcomes. The sensor in Bengaluru shows no change in its relationship to the BAM over one year. Meanwhile the sensor in Delhi reads lower compared to the BAM after a year. Possible reasons for the change in Delhi include a loss of sensitivity due to contamination inside the chamber or aging of components. Other explanations could be differences in aerosol size distribution or chemical composition between the two winters changes in the reference instrument or how its data is processed or a replacement of the sensor unit. The
public archive does not keep records of sensor identity, maintenance or calibration history. So it is not possible to separate these factors. Checks on concentration bands and humidity rule out the simplest possibilitieslike a change in pollution levels or relative humidity.
Practical implications. Both the season and the passage of time affect sensor error, so a calibration should be tied to the season it was made in. In high-concentration urban sites such as Delhi, it should be checked against a reference at least once a year and not assumed to hold beyond that. In our Delhi data, a calibration fitted in one
year lost much of its accuracy a year later. At lower- concentration sites such as Bengaluru, we found no detectable change over a year, so a single calibration may stay usable for longer, although the relative error there is larger. Reporting alongside bias, and comparing matched seasons, helps avoid mistaking the seasonal cycle for sensor ageing.
Value of the instrument. The virtual instrument provides a inspectable process where each statistical step appears as a visible block and every result is shown graphically. Its outputs can reproduce an implementation showing the precision of the results. Since it works with files it can be expanded to collect data from a sensor via a DAQ or serial connection. It also flags when bias or moves away from a baseline that’s specific, to the season.
Limitations. This study covers only two sites and at most two annual cycles. The season-matched windows are short: two months for Bengaluru and three for Delhi. The archive gives one series per site, with no information on which sensor unit was used or whether it was replaced. Only the CF1 channel was analysed. The BAM reference has its own measurement uncertainty. Finally, the VI currently processes only the Bengaluru file, and the Delhi results come from Python. The bootstrap method accounts for within-site autocorrelation. It does not handle unmeasured differences, between years. The results should be treated as a case study; they should be replicated with records and multiple units ideally with sensor-level metadata.
-
CONCLUSION
A LabVIEW virtual instrument was built to calculate performance metrics and trend indicators for collocated low-cost PM2.5 sensor data. The results were checked against code to ensure accuracy. When applied to PurpleAir PA-II and BAM data the analysis showed that sensor error varies greatly with the seasonbias in Bengaluru ranged from 0 to +24 g m3. A simple monthly trend of this error (+0.35 g m3 per month) does not represent a drift.
A comparison matched by season found no drift in Bengaluru ( = +0.02, 95 % CI 0.06 to +0.11). In Delhi however there was an significant drop in the sensor-to-reference ratio ( = 0.35, 95 % CI 0.40 to
0.29). This change reduced the accuracy of a calibration made a year earlier increasing the NRMSE from 15 % to 33 %.
We recommend that evaluations of low-cost sensors always compare data from the seasons before assuming that changes are due to sensor ageing. In cities in India, with pollution levels it is essential to carry out regular reference checks and recalibrate sensors periodically.
REFERENCES
-
M. J. Campmier et al., Seasonally optimized calibrations improve low-cost sensor performance: long-term field evaluation of PurpleAir sensors in urban and rural India, Atmos. Meas. Tech., vol. 16, no. 19, pp. 43574374, 2023, doi: 10.5194/amt-16-4357-2023.
-
K. K. Barkjohn, B. Gantt, and A. L. Clements, Development and application of a United States-wide correction for PM2.5 data collected with the PurpleAir sensor, Atmos. Meas. Tech., vol. 14, no. 6, pp. 46174637, 2021, doi: 10.5194/amt-14-4617-
2021.
-
T. Sayahi, A. Butterfield, and K. E. Kelly, Long-term field evaluation of the Plantower PMS low-cost particulate matter sensors, Environ. Pollut., vol. 245, pp. 932940, 2019.
-
F. M. J. Bulot et al., Long-term field comparison of multiple low-cost particulate matter sensors in an outdoor urban environment, Sci. Rep., vol. 9, art. 7497, 2019.
-
C. Malings et al., Fine particle mass monitoring with low-cost sensors: Corrections and long-term performance evaluation, Aerosol Sci. Technol., vol. 54, no. 2, pp. 160174, 2020.
-
M. J. Campmier et al., Seasonally optimized calibrations improve low-cost sensor performance: Long-term field evaluation of PurpleAir sensors in urban and rural India, Dryad dataset, 2023, doi: 10.6078/D1RQ70.
-
H. B. Mann, Nonparametric tests against trend, Econometrica, vol. 13, no. 3, pp. 245259, 1945.
-
M. G. Kendall, Rank Correlation Methods, 4th ed. London, U.K.: Griffin, 1975.
-
P. K. Sen, Estimates of the regression coefficient based on Kendalls tau, J. Amer. Statist. Assoc., vol. 63, no. 324, pp. 13791389, 1968.
-
B. Efron and R. J. Tibshirani, An Introduction to the Bootstrap. New York, NY, USA: Chapman & Hall, 1993.
-
W. K. Newey and K. D. West, A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix, Econometrica, vol. 55, no. 3, pp. 703708, 1987.
