Large data sets comprising diagnoses about chronic conditions are becoming increasingly available for research purposes. In Germany, it is planned that aggregated claims data including medical diagnoses from the statutory health insurance with roughly 70 million insurants will be published on a regular basis. Validity of the diagnoses in such big data sets can hardly be assessed. In case the data set comprises prevalence, incidence and mortality, it is possible to estimate the proportion of false positive diagnoses using mathematical relations from the illness-death model. We apply the method ...