A strong laboratory score can tell a formulation team that an e-liquid is clean, recognisable and well balanced under controlled conditions. It cannot, by itself, predict whether adult consumers will prefer that liquid across different devices or keep choosing it over time.
That distinction matters because e-liquid consumer preferences are shaped by more than aroma complexity. Sweetness, cooling, harshness, nicotine delivery, hardware settings and consistency from one batch to the next can all change the experience.
The useful question is therefore not whether the laboratory or the consumer is “right.” It is whether a development programme measures technical quality, sensory liking and real-device performance as separate but connected outcomes.
Lab quality and consumer liking measure different things
Laboratory and trained-panel work is valuable for identifying off-notes, checking flavour identity, comparing intensity and assessing whether a formula meets its specification. Consumer testing asks a different question: does the whole experience feel pleasant enough to use again?
A controlled psychophysical study of six commercial e-cigarette flavours found that liking was positively associated with perceived sweetness and cooling, while harshness was negatively associated with liking for several flavours. The study was small and does not forecast market sales, but it shows why a technically complex flavour is not automatically the most appealing one.
In other words, more layers are not the same as a better experience. A restrained formula can perform well if its sweetness, cooling, bitterness and harshness sit in a range that the intended adult user finds comfortable.
Device power and liquid composition change the delivered experience
An e-liquid is not experienced in isolation. The device, coil, power output and puff conditions determine how the liquid is heated and how much aerosol reaches the user.
In one laboratory study, researchers varied propylene glycol and vegetable glycerin composition across three device-power settings. They found that the effect of PG/VG composition on nicotine emissions depended on power, and that higher power reduced some of the differences between the tested liquids. Other emission studies have likewise found that device design and liquid characteristics interact.
That does not mean every formula must work identically in every device. It means the target hardware and operating range should be part of the formula brief from the beginning. A flavour approved only on one reference setup may need further validation before it is used in another pod, coil or power range.
For a closer look at laboratory methods, see VAPEAST’s guide to e-liquid quality control beyond GC-MS.
Why complexity can work against repeat comfort
Adding more flavour components can create depth, but it also gives the development team more interactions to control. Each additional note must remain recognisable after storage, heating and dilution into the finished aerosol.
The practical risk is not that complex formulas are inherently poor. It is that complexity can be mistaken for premium quality before the blend has been tested for comfort and stability. An impressive first draw may still become tiring if sweetness, cooling or intensity accumulates during longer use.
A better development target is durable clarity: the intended flavour remains recognisable, harshness stays controlled and the experience does not change sharply across the supported device range.
Engineering turns a formula into a repeatable product
Once a formula moves beyond the bench, product quality depends on more than the flavour recipe. Raw-material controls, mixing and conditioning, filling accuracy, storage, coil compatibility and batch-release testing all affect what reaches the consumer.
That is why a reliable programme should separate five questions:
| Evaluation layer | What it should answer | Useful evidence |
|---|---|---|
| Formula identity | Does the blend match its intended profile? | Analytical checks and trained-panel descriptors |
| Hardware compatibility | Does it perform across the intended coil and power range? | Standardised puffing and device comparison |
| Sensory comfort | How do sweetness, cooling, bitterness and harshness affect liking? | Blinded adult-user sensory testing |
| Production consistency | Do batches stay within specification? | Incoming-material, in-process and release controls |
| Market durability | Do complaints, returns and repeat purchases support the product? | Post-launch data tracked separately from lab scores |
A practical test plan for product teams
- Define the target hardware. Record coil type, resistance, power range, airflow and intended nicotine strength before final flavour approval.
- Use controlled sensory methods. Blind samples, randomise order and score specific attributes as well as overall liking.
- Test more than the first draw. Include repeated-puff sessions so sweetness, cooling, harshness and flavour fatigue can be assessed.
- Run stability and batch checks. Compare fresh, stored and production-batch samples against an approved reference.
- Keep market metrics separate. Complaints, returns and repeat purchasing can validate a launch, but they should not be presented as laboratory findings.
What the evidence does—and does not—show
Published studies support two grounded conclusions: sensory attributes influence flavour liking, and device power can interact with liquid composition to change emissions. They do not prove that a high blind-test score causes a lower repurchase rate.
That stronger claim should be treated as an industry hypothesis unless a company publishes its test design and sales data. The responsible editorial takeaway is narrower: no single laboratory score captures the full consumer experience.
Final take
The best formula is not simply the most complex sample or the one with the highest score on a single tasting day. It is the formula that meets its technical specification, stays consistent in its intended hardware and performs well in appropriately designed adult-user testing.
R&D raises the ceiling, but a repeatable product also needs a well-defined floor. When sensory, hardware and production evidence are reviewed together, teams can make better decisions than they can from either laboratory judgement or market anecdotes alone.






