Why dose context matters: fragrance hazard vs. real‑world risk — lessons from benzaldehyde and p‑cymene
Table of Contents
- Key Highlights
- Introduction
- Clarifying the difference: hazard versus risk
- Why benzaldehyde and p‑cymene were chosen
- Converting abstract toxicology into intuitive examples
- How aggregate exposure modeling anchored the analysis
- Conservative assumptions: intentionally overestimating exposure
- Margin of Safety: how to interpret the numbers
- Why RIFM has shifted away from new animal testing
- Practical implications for product developers and formulators
- Communication challenges: hazard language and consumer perceptions
- Limitations, uncertainties, and what the data do not address
- What the findings mean for regulation and standards
- Next research steps and the path forward
- Practical guidance for stakeholders
- A final perspective on evidence and public discourse
- FAQ
Key Highlights
- Toxicology identifies hazards using high experimental doses; realistic exposure data and aggregate modeling show those hazard-level doses are far above what consumers encounter with fragrances.
- Case studies on benzaldehyde and p‑cymene show estimated annual consumer exposure is measured in hundredths of a milliliter — producing margins of safety in the tens to hundreds of thousands when compared with conservative toxicological effect levels.
Introduction
Debates over ingredient safety often pivot on two different scientific questions: what a substance can do in a controlled test and what it will actually do when people use products as intended. Toxicology delivers the first answer by detecting hazards at targeted, frequently high doses. Risk assessment supplies the second by combining those hazard findings with realistic exposure estimates to determine whether a real‑world threat exists.
A recent paper by Kaushal Joshi, PhD, DABT, Arianna Bartlett, PhD, and colleagues from the Research Institute for Fragrance Materials (RIFM) confronts this distinction head‑on. Using benzaldehyde and p‑cymene as illustrative case studies, the authors compare toxicological effect levels from animal studies with probabilistic, aggregate consumer exposure calculations. The result places hazard findings in a practical frame: typical fragrance use delivers only trace amounts of these ingredients, generating safety margins far beyond conservative regulatory thresholds.
The arguments and data in the paper carry implications for industry safety assessments, product development, regulatory communication, and consumer understanding. What follows is a detailed unpacking of the study’s methods, findings, assumptions, and implications — together with concrete examples and practical guidance for stakeholders who must interpret hazard information alongside exposure science.
Clarifying the difference: hazard versus risk
Hazard describes what a substance is capable of causing under specified conditions, frequently identified through high‑dose laboratory studies. Risk combines that inherent potential with how much, how often, and by which route people are actually exposed.
Toxicological tests are designed to detect hazards efficiently. They often escalate doses until an effect appears, or they use a range that ensures a clear dose–response relationship. Those doses can be orders of magnitude higher than anything a consumer would experience during normal product use. That design is scientifically sound for hazard detection but can mislead when hazard data are presented without exposure context.
Risk assessment fills that gap. It translates a NOAEL (no‑observed‑adverse‑effect level) or a lowest‑observed‑adverse‑effect level into a safe exposure point for humans by applying uncertainty factors. It then compares that safe point with realistic exposure estimates, which should reflect multiple products, routes of entry (dermal, inhalation, oral), and user behaviors. When that comparison yields a margin of safety (MOS) above accepted benchmarks, the conclusion is that normal use does not present a consumer health concern.
This distinction has practical consequences. Hazard‑only narratives are deterministic and can fuel alarm when removed from dose context. Exposure‑driven risk assessments provide a probabilistic, practical perspective that informs decisions about formulation, use limits, labeling, and communication.
Why benzaldehyde and p‑cymene were chosen
Selecting illustrative compounds requires balancing scientific rigor with accessibility. Benzaldehyde and p‑cymene meet both criteria.
Both are recognized fragrance ingredients with publicly available in vivo toxicology data, enabling identification of conservative NOAELs and adverse‑effect levels. Both occur naturally in foods: benzaldehyde is a key component of almond flavor and present in many fruits; p‑cymene exists in citrus and other plant oils. That overlap with common foods allows for everyday analogies that clarify the magnitude of the exposure gap between toxicology studies and normal human experience.
Using these two molecules, the authors demonstrate a general principle rather than making narrow claims about every fragrance. The paper shows how hazard endpoints become practically irrelevant when compared with typical human exposure for many fragrance constituents — and how aggregate exposure modeling provides the appropriate context for safety decisions.
Converting abstract toxicology into intuitive examples
Toxicology typically reports doses in mg/kg body weight per day. Those units are necessary for cross‑species comparisons and regulatory calculations, but they are abstract for non‑specialists. The paper converts those numbers into concrete examples to make the scale more tangible.
For benzaldehyde, the authors calculated that reaching the conservative safe dose inferred from toxicology would require consuming roughly 83,000 almonds per day for a lifetime. For p‑cymene, the comparable food example was about 153,778 raspberries per day for a lifetime.
Fragrance use analogies were equally striking. Using conservative assumptions, the study estimated that an individual would need to apply approximately 138,330 perfume sprays per day (for life) containing benzaldehyde, or about 18,870 sprays per day (for life) containing p‑cymene, to approach the conservative toxicological thresholds. Those scenarios are patently unrealistic. That point is the purpose of the exercise: to make the magnitude of the gap between experimental hazard doses and real‑world exposure immediately visible.
Presenting results as everyday analogies does not replace formal risk characterization; it complements it. Analogies help non‑technical audiences grasp why a statistically significant effect in a high‑dose animal study does not automatically translate into a human safety hazard under normal product use.
How aggregate exposure modeling anchored the analysis
The exposure side of the equation rested on the Creme RIFM Aggregate Exposure Model, a probabilistic tool that integrates extensive consumer use and behavior data across geographies and product categories. The model calculates individual and aggregate exposure across oral, dermal, and inhalation routes and aggregates exposure from cosmetics, personal care, household products, and air care items.
Key methodological choices:
- Use of large consumer datasets from North America, Europe, and Asia to capture variation in product use patterns.
- Probabilistic estimation to produce distributions of exposure, allowing assessment of median and high‑end (for example, 95th percentile) users.
- Aggregation across product types and routes to approximate real‑world cumulative exposures.
For benzaldehyde, the model’s total chronic aggregate exposure estimate was 0.00053 mg/kg body weight/day. For p‑cymene, it was 0.00061 mg/kg body weight/day. Those are average chronic exposures reflecting high‑end use scenarios as represented in the probabilistic distributions.
Translating those daily per‑kilogram figures into physical quantities strengthens the intuitive grasp of scale. For a 60 kg adult — the conservative body weight used in the calculations — the daily exposures are about 0.0318 mg/day for benzaldehyde and 0.0366 mg/day for p‑cymene. Over a year, those daily amounts sum to roughly 11.6 mg (0.0116 g, ~0.011 mL) and 13.4 mg (0.0134 g, ~0.013 mL), respectively — fractions of a single milliliter across an entire year of product use.
Those volumes are equivalent to about 0.2 and 0.3 drops per year. Presenting exposure as drops-per-year is a vivid counterpoint to the milligrams‑per‑kilogram numbers that dominate toxicology reports.
Conservative assumptions: intentionally overestimating exposure
The study made intentionally conservative choices to avoid underestimating exposure. Conservative assumptions increase confidence that safety conclusions are not artifacts of optimistic modeling.
Examples of conservative assumptions used:
- 100% absorption across the relevant exposure route. Real dermal and inhalation absorption are typically lower than 100% for many fragrance ingredients, so this assumption intentionally inflates modeled systemic uptake.
- No evaporation following product application. In practice, a portion of volatiles dissipates rapidly, reducing systemic exposure.
- Use of high‑end values for ingredient concentrations drawn from food examples and product surveys when higher variability could have been modeled with narrower central tendencies.
- A default body weight of 60 kg, which is lower than average adult body weight in several high‑income countries, producing a more conservative per‑kilogram exposure estimate for a given absolute intake.
These conservative assumptions bias exposure estimates upward. The MOS and analogue calculations therefore err on the side of caution. The continued finding of extremely large safety margins under those conditions strengthens the central conclusion: real‑world exposures are minute relative to toxicology effect levels.
Margin of Safety: how to interpret the numbers
The margin of safety (MOS) is the ratio of an established safe dose (often derived from a NOAEL in animal studies, adjusted by uncertainty factors) to the estimated human exposure. Regulators commonly accept an MOS of 100 or greater as an indicator of low concern for general human populations. That benchmark stems from standard uncertainty factors: 10 for interspecies differences (animal to human) and 10 for intraspecies variability (human-to-human differences), yielding a combined factor of 100.
The paper reported MOS values of:
- 377,358 for benzaldehyde
- 81,967 for p‑cymene
Those figures exceed the conventional safety benchmark by several orders of magnitude. The MOS values corroborate the food and perfume spray analogies: even under conservative exposure assumptions, the gap between detected hazard doses and realistic human exposure is extremely large.
A high MOS does not imply absolute zero risk for all possible endpoints. It indicates that, based on the available toxicological data and the modeled exposures, there is a substantial buffer between typical exposure and doses associated with adverse findings in experimental studies. The higher the MOS, the greater the confidence that normal use poses negligible systemic risk to the general population.
Why RIFM has shifted away from new animal testing
RIFM’s approach reflects a broader evolution in toxicology and regulatory science. The organization has not commissioned new animal toxicity studies for human health endpoints in more than a decade. Instead, RIFM emphasizes:
- Reuse and integration of existing, robust toxicological datasets.
- Advanced exposure modeling to place those datasets into human‑relevant contexts.
- Targeted use of non‑animal methods where data gaps justify additional investigation.
Aggregate exposure modeling and probabilistic approaches enable meaningful risk assessments without routine reliance on new high‑dose animal studies. This approach aligns with the 3Rs principles (replace, reduce, refine) and with the regulatory trend toward weight‑of‑evidence assessments that integrate multiple streams of data — in vivo, in vitro, and human observational studies — together with exposure science.
The shift also reflects greater availability of human use data: thousands of consumers have been surveyed about product use frequency, amounts, and combinations. Those data improve exposure estimates and make animal testing less necessary for routine safety confirmation when existing toxicology is high quality.
Practical implications for product developers and formulators
Product teams must navigate two responsibilities simultaneously: ensure formulations meet regulatory safety criteria and communicate ingredient safety credibly to consumers. The RIFM analysis yields several operational takeaways.
-
Build exposure‑aware formulations. Safety evaluation should begin with realistic exposure modeling. Limiting an ingredient based solely on hazard reports without considering actual exposure may unnecessarily constrain formulation options or lead to formulations that are no safer but are less effective or desirable to users.
-
Prioritize high‑quality data over headline‑driven signals. Toxicological findings from high‑dose animal studies should trigger a focused data evaluation rather than immediate reformulation. Cross‑referencing hazard endpoints with aggregate exposure and MOS calculations can clarify whether a hazard finding has practical relevance.
-
Consider cumulative and aggregate exposure across a brand’s product portfolio. Many consumers use multiple products from the same brand. Aggregate modeling can guide concentration decisions across products to keep cumulative exposure well within conservative safety margins.
-
Use conservative assumptions when developing internal safety cases for marketing claims or voluntary limits. The RIFM approach illustrates how conservative modeling reduces the chance of underestimating exposure and results in robust safety margins that support confident product decisions.
-
Be transparent and contextual in communication. Hazard statements without exposure context can sow confusion. Providing clear, accessible explanations that pair hazard information with exposure and MOS helps consumers and regulators understand the real‑world implications.
Communication challenges: hazard language and consumer perceptions
Public conversations about ingredient safety often emphasize hazard signals. Advocacy campaigns, media reporting, and simplified labeling can spotlight chemical names or isolated study outcomes without the exposure context that determines actual risk.
That dynamic creates a persistent challenge for the personal care sector: how to balance transparency and responsiveness with scientific accuracy. The case studies in this paper suggest practical approaches:
- Present hazard findings alongside quantified exposure estimates and MOS values. That pairing enables consumers to understand both the potential for harm in principle and the likelihood of harm in practice.
- Use analogies sparingly and accurately. Everyday comparisons — almonds, raspberries, drops of fragrance per year — can be powerful, but they must be calibrated to avoid trivializing concerns for sensitive subpopulations or for endpoints that require a different analytical framework (for example, sensitization).
- Distinguish systemic toxicity from local effects. Fragrance materials can cause contact allergy in sensitized individuals, which is a different risk mode from systemic toxicity assessed via NOAEL/MOS calculations. Clear delineation helps consumers and clinicians make informed choices.
Balanced communication reduces the chance that hazard labels prompt unnecessary substitution of ingredients that pose no actual consumer risk while potentially increasing environmental or functional liabilities.
Limitations, uncertainties, and what the data do not address
Robust as the analysis is, it does not eliminate uncertainty entirely. Key limitations and points of caution:
- MOS focuses on systemic endpoints derived from available toxicology. It does not directly address allergic contact dermatitis or sensitization potential, which depend on exposure patterns, skin chemistry, and individual predisposition.
- The aggregate model relies on available consumer use data. Although those datasets are large and geographically diverse, some niche product uses or emerging product formats may not be fully represented.
- Interspecies differences in toxicity mechanisms can complicate direct extrapolation of some endpoints. The paper’s next steps include investigating biological relevance across species and targeted in vitro assays to clarify mechanisms where needed.
- The conservative assumptions — while deliberately protective — may obscure the precise, likely lower exposures in real use. The conservative bias is intentional, but it means the modeled values should be interpreted as upper‑bound, not central‑estimate, exposure scenarios.
Acknowledging these limitations is not a reason to disregard the conclusions. It is a call for continuous refinement: integrating better human biomonitoring data where available, expanding consumer datasets for new product formats, and using modern in vitro and computational methods to bridge species differences.
What the findings mean for regulation and standards
Regulatory bodies apply uncertainty factors and conservative assumptions routinely. The RIFM analysis speaks directly to how regulators should interpret hazard data: without context, alarms can arise from unavoidable aspects of hazard identification testing.
Specific regulatory implications:
- Regulators and standard‑setting organizations should continue to require exposure assessment as an integral part of ingredient safety evaluation. The MOS threshold (commonly 100) remains a practical benchmark, but calculating it responsibly requires realistic aggregate exposure estimates.
- Harmonizing exposure modeling approaches improves comparability of safety decisions. Organizations and regulators using different assumptions can produce widely divergent exposure estimates. Greater alignment on conservative assumptions, data inputs, and probabilistic methods will strengthen regulatory coherence.
- Where animal data are old or incomplete, targeted in vitro tests or modern computational toxicology can be used to address mechanism‑specific questions instead of wholesale new high‑dose animal studies.
The combination of robust hazard data, probabilistic aggregate exposure modeling, and targeted mechanistic investigations produces a defensible, transparent basis for regulatory decisions that protect consumers without imposing unnecessary constraints on formulations.
Next research steps and the path forward
The paper points toward several practical follow‑ons:
- Evaluate additional fragrance ingredients using the same framework to determine whether the exposure‑vs‑hazard separation observed for benzaldehyde and p‑cymene holds broadly.
- Investigate the mechanistic relevance of high‑dose animal effects for human biology. Some effects observed at high doses may involve pathways that are not relevant at low exposures or that are species‑specific.
- Conduct targeted in vitro assays to fill critical data gaps and strengthen weight‑of‑evidence assessments. These assays can clarify whether an observed hazard has plausible relevance at human exposures.
- Expand consumer use databases to capture new products, evolving formulations, and changing use patterns. As formats such as sprays, aerosols, and concentrated consumer products evolve, exposure models must adapt accordingly.
- Improve biomonitoring where possible. Human biomonitoring data provide powerful direct evidence of systemic exposure and can validate modeled estimates.
Collectively, these steps narrow uncertainty, reduce reliance on broad extrapolations, and support a mature, exposure‑aware approach to safety.
Practical guidance for stakeholders
For industry leaders:
- Embed aggregate exposure modeling early in product design and ingredient selection.
- Maintain high‑quality records of product composition and consumer use to support probabilistic modeling.
- Adopt conservative assumptions for internal safety cases, but communicate the resulting MOS and assumptions clearly.
For regulators:
- Require exposure context when evaluating hazard reports submitted for ingredient assessments.
- Encourage harmonization of modeling approaches across jurisdictions.
- Promote use of non‑animal methods for mechanism‑specific questions.
For clinicians and consumer advocates:
- Recognize the difference between hazard reports and realistic risk for systemic toxicity.
- Continue to monitor and communicate local allergic reactions and sensitization risks, which require a different assessment framework than MOS calculations for systemic endpoints.
- Advocate for transparent, accessible communication that includes exposure context and MOS for ingredients of public concern.
For consumers:
- Recognize that a chemical’s detection in a high‑dose animal study does not automatically mean everyday products pose a risk.
- If you have a history of skin sensitivity or allergy, consult product ingredient lists and patch testing guidance rather than relying on hazard headlines.
A final perspective on evidence and public discourse
The study by Joshi, Bartlett, and colleagues reasserts a fundamental toxicological truth: dose matters. Hazard identification is the first essential step. Without realistic exposure assessment, hazard findings can create misplaced concern.
Fragrance safety assessments that combine well‑characterized hazard data with advanced aggregate exposure modeling deliver practical answers. For benzaldehyde and p‑cymene, the modeled exposures translate into fractions of a milliliter per year and margins of safety that exceed regulatory benchmarks by large factors. Those findings do not trivialize the importance of vigilance where sensitization or niche high‑exposure scenarios exist. They do, however, provide a clearer basis for regulatory decisions, formulation choices, and public communication.
The approach demonstrated in this paper points toward a science‑driven middle ground: one that honors rigorous hazard detection, applies conservative exposure modeling, and uses transparent metrics — like MOS — so that stakeholders can make informed choices grounded in both toxicology and realistic human use.
FAQ
Q: What is the difference between hazard and risk? A: Hazard is the inherent ability of a substance to produce an adverse effect under specified conditions, typically identified in controlled tests. Risk combines hazard with exposure — it answers whether people will encounter the substance at levels and by routes that make that hazard relevant.
Q: Why do toxicology studies use high doses? A: High doses increase the likelihood of observing an effect and help establish dose–response relationships. That design is efficient for hazard identification, but those doses are often far above consumer exposures, so translating findings into risk requires exposure data.
Q: How did the authors estimate consumer exposure to benzaldehyde and p‑cymene? A: They used the Creme RIFM Aggregate Exposure Model, a probabilistic tool that integrates large consumer use datasets across geographies and product categories and aggregates exposure across dermal, inhalation, and oral routes. The model generated chronic per‑kilogram‑per‑day exposure estimates, including high‑end users.
Q: What conservative assumptions were used in the exposure modeling? A: Examples include assuming 100% systemic absorption, assuming no evaporation after application, using high‑end ingredient concentrations where available, and applying a conservative adult body weight of 60 kg. These assumptions bias exposure estimates upward to avoid underestimation.
Q: What does a margin of safety (MOS) mean, and what is considered safe? A: MOS is the ratio of a safe dose (often derived from a NOAEL with uncertainty factors applied) to estimated human exposure. Regulators commonly consider an MOS of 100 or greater as protective for general populations because it includes standard uncertainty factors for interspecies and intraspecies variability.
Q: What MOS values did the study find for benzaldehyde and p‑cymene? A: For benzaldehyde, the MOS was 377,358. For p‑cymene, the MOS was 81,967. Both values exceed the commonly used conservative benchmark of 100 by large margins.
Q: Do these findings mean fragrances are completely risk‑free? A: No single analysis eliminates every potential concern. The findings indicate that systemic risks from normal use of these specific fragrance ingredients are extremely low when exposure is realistically modeled. Other endpoints, such as allergic contact dermatitis in sensitized individuals, require separate assessment methods. Niche scenarios with unusually high exposure could also demand focused evaluation.
Q: Why has RIFM not conducted new animal human health studies for more than a decade? A: RIFM has shifted toward integrating existing robust toxicology datasets with advanced exposure modeling and targeted in vitro or computational methods where needed. This approach reduces reliance on new animal testing while preserving high standards for safety assessment.
Q: How should manufacturers use this information when formulating products? A: Incorporate aggregate exposure modeling into safety evaluations early; prioritize data quality over sensational hazard narratives; maintain conservative internal assumptions for safety cases; and communicate hazard and exposure context transparently to consumers and regulators.
Q: What are next steps in this line of research? A: The authors plan to evaluate more fragrance ingredients using the same exposure‑anchored framework, investigate the biological relevance of high‑dose animal effects to human biology, and use targeted in vitro assays to address specific data gaps.
Q: How can consumers interpret media reports that highlight hazard findings? A: Look for context: does the report include exposure estimates or a margin of safety? If it mentions only high‑dose animal findings without exposure context, the practical implications for everyday product use may be limited. For personal skin sensitivity concerns, consult a clinician or patch testing guidance.
