Data

Sample

Item Value
Raw responses 282
Dropped for missing data 38
Analytic sample 244
Distinct firms 219
Firms appearing in more than one wave 48 (approximately 22%)
Observations from multi-wave firms 111
Firms appearing once 171
Waves 2022–2025
Firm size range 1–300 employees

Missingness was tested before deletion rather than after; the test was consistent with data missing completely at random, which is what licenses listwise deletion here.

Composition

Wave N
2022 32
2023 71
2024 73
2025 68

Sector composition in the first wave: manufacturing 56%, IT 22%, services 22%.

Geographic distribution: Capital area 38%, Yeongnam 31%, Chungcheong 18%, Honam 13%.

Measures

Construct Operationalization
Technology (T) Digital transformation awareness, self-reported
Environment (E) Smart-system infrastructure score
Organization (O) Firm size, log-transformed
Training exposure Programme participation intensity
Outcome Post-minus-pre DT competency, single five-point self-report item

Test–retest reliability on the multi-wave subset: T r=.78, E r=.82.

ImportantMeasurement problem under revision

The outcome is a single self-report item answered by the same respondent who supplies the T and E predictors. The study’s own argument is that single self-report items under-capture training effects — which means the instrument that motivates the argument is also the instrument the estimates depend on. Resolving this is the first item of the pending revision, and it changes the dependent variable rather than merely adding a robustness check.

Structural caveat

The distinction between 219 firms and 244 observations matters for method selection. Methods that require within-firm variation are identified only on the 48 firms observed more than once. Applying them to the pooled sample and relying on year fixed effects does not recover firm-level identification.