- Sensitive company or customer material is uploaded to an unapproved workspace
- Classify the material first, use an approved enterprise environment, minimize the files, and remove anything outside the workspace policy.
- The scraper violates platform terms, exceeds rate limits, or stores unnecessary user data
- Use the supported API or licensed export, narrow the query, honor limits and deletion requirements, and collect only fields needed for research.
- Reddit or another public forum is described as unbiased or representative of all customers
- Document source bias and coverage, compare with other research, and present the findings as one signal rather than market truth.
- The model invents frequencies or produces quotes that cannot be found in the dataset
- Calculate counts deterministically, require row or URL citations, spot-check each top theme, and retain counterexamples.
- New conclusions accumulate without dates, versions, or a canonical artifact
- Version raw data and analysis, add reviewed findings to a dated artifact, and replace superseded material deliberately.