It’s nearly an unquestioned truth these days that the best product organizations thrive on experimentation. Recent breakthroughs in LLMs and generative AI are another reminder of the importance of utilizing experiment-driven thinking and processes (as embodied by our Continuous Hackathon (opens in new tab)) to quickly create products that have verifiable positive impacts for end users and the business bottom line. As a relative newcomer to Oscar, I embarked on an exploration of how Oscarians think about and use experimentation in their work in the spirit of our Seek Insight product principle.
Oscar Product Principle: Seek Insight
Why experimentation matters
A few themes emerged when I asked my colleagues about why they experiment, and what they get out of experimentation.
Verify the impact of a change: As our product principle above states, we can and should assert our views about user psychology. Product sense/intuition is often a critical input to the product development process. Yet, experimentation lets us balance that intuition with hard facts, and grounds us in observed user behavior that may be quite different from what we hypothesized. Tony Evans, a PM and former Product Analytics team member, points out that we can’t solely rely on our intuition because we are building products for members that are often quite different from us demographically and in terms of lived experience.
We are engrossed in insurance but our members are not, so messaging we think might resonate with members sometimes doesn’t. Danielle Ampulski on our Product Marketing team cites an instance where we thought that messaging that highlighted benefits like virtual urgent care would be more effective at encouraging members to create digital accounts. As it turned out, a more transactional reminder that didn’t tout the benefits of account creation actually did better!
Enable ramp ups, rollbacks, and pivots: Experiments are often set up in such a way that they can enable us to carefully manage the scope and impact of a change. Common tools used in classic A/B split testing such as feature flags allow us to gradually ramp changes up to a wider share of users, or rollback changes entirely if an observed change doesn’t meet our success criteria. Brian Lee, a PM for one of our teams focusing on virtual care, cites the ability to rollback if a bug arises as another benefit of feature flags.
Emili Hsu, a Senior Designer, recounts an experiment where we changed a step in our account creation flow where we ask for members’ personal information. We added fields for middle name and suffix, and defaulted to the most commonly used ID verification method–social security number. Our hypothesis was that these changes would improve information accuracy. However, our experiment results showed that form completion rates worsened, likely due to the form appearing longer and more overwhelming upfront, and because of user sensitivities around being asked for their social security number so explicitly. Our experiment set up let us easily rollback to the control experience while we digested learnings from the experiment and rethought our approach.
Side-by-side comparison of the experiment V.S. control member information form
Capturing unexpected outcomes: Experiments that verify the impact of a change or a key hypothesis are ideal scenarios. Sometimes, it’s the unexpected outcomes of an experiment that grab your attention. Jaime Morrison on our Product Marketing team cites a recent example where starting with a smaller pilot group for an email campaign where they were experimenting with messaging and incentives to encourage members to take advantage of home infusion benefits uncovered that members were made uneasy by the unexpected monetary compensation. This early insight allowed her to make further adjustments to the campaign before ramping up the experiment.
What we experiment on
High volume member facing applications/pages: These are the most obvious targets for experimentation because we can readily obtain statistically significant data on them. Herry Pierre-Louis, a Senior PM for some of our member-facing teams, notes that even small changes to these experiences can create multiple downstream impacts, making experimentation even more prudent. As an example, we found that reordering a single question to an online booking process for virtual care, combined with adding a search icon and showing suggestions, resulted in a 17% improvement in conversion for appointment bookings.
Lower volume internal tools: The category above may be the ripest for experimentation, but Erin Landau, a Group PM, advocates for experimenting with lower volume internal tools as well. We sometimes assume that a retrospective evaluation of changes to an internal tool is sufficient, but it’s hard to attribute the impact of a change retroactively. Many factors may have changed concurrently alongside the change to the product itself, which are difficult to control for after the fact.
One example of internal tools experimentation is an ongoing A/B test that’s assessing the impact of our new Pharmacy claims and benefits aggregation tool. Previously, our pharmacists had to log into several different applications to understand what drugs were in network and gauge a member’s utilization of specific medications. We built an application to centralize all this information for our pharmacists, and pitted our new app in a head to head test against the old process. While the test is still underway, early results indicate a significant decrease in the time it takes for our pharmacists to complete their daily work.
Internal processes: Experimentation at Oscar isn’t limited to classic A/B split tests on digital products, it extends to processes and programs as well. Josh Vallon, Senior Software Engineer for Engineering Enablement, discussed experimenting with different ways to create and maintain a single source of truth for cross-functional team directories (basically every product pod by definition!). This type of experimentation touches on process, change management, tooling, and automation. Meanwhile, he also mentions experimenting on programs such as making changes to how we conduct mentoring on our Engineering teams. These types of experiments can be especially challenging because of the long cycle times required to see results, as well as the qualitative nature of the outcomes.
Product marketing emails: Our teams take full advantage of our homegrown engagement and automation platform Campaign Builder to run experiments that tweak the messaging of campaigns as well as provide different incentives to drive member actions. Danielle and Jaime explain that they experiment both on one-off campaigns as well as always-on, recurring series of emails. An example of a one-off campaign experiment is the one discussed earlier where Jaime is exploring whether incentives can influence members to choose home infusions. Danielle discusses that for always-on campaigns, they will assess the performance of each email in the series, and will try to replace low-performing emails with new variants. The Product Marketing team roadmaps the core hypotheses for these always-on campaigns that they want to test for the full year, so that they know what resources they may need well in advance.
How we go about experimentation
Let’s take a look at the approach to experimentation typically used by Herry for member experience products as an illustrative example of how we practice classic A/B experimentation on Oscar’s Product team.
When devising an experiment, we start with the company mission and vitals, and zero in on those relevant to our work. From here, we can develop strong hypotheses about what outcomes we’re trying to achieve with the changes we’re rolling out. Ideally, previous user research (both quantitative and qualitative), surveys, and observational studies will have informed and refined our experiment design and hypotheses, such that the experiment itself serves to provide ground truth confirming this prior research. We can think about KPIs impacted based on the hypotheses formulated, and develop success metrics subsequently.
Now that we have a hypothesis and success metrics in hand, we can collaborate with our data science colleagues to project the time to significance based on projected volume (be it clicks, impressions, etc.). The length of an experiment will ultimately depend on the volume of interactions for that given product, what kind of signal we’re seeking, and whether it needs to be statistically significant. We inform our internal stakeholders about experiments in progress, so that they’re aware there may be upcoming changes for members. We monitor the experiment weekly, and make a go/no-go decision at the end about whether to ramp experiments up to 100%. If we decide to rollout in full, we will inform our business stakeholders as such and publish an experiment wrap-up.
Best practices for experimentation
Be intelligent about what metrics to track because this is key to successful experimentation. Alan Chen, one of our Data Scientists, recommends picking metrics that are most directly attributable to the change you’re introducing. Ideally, these metrics can roll up into higher level metrics that matter to the business. An example of this would be if we were experimenting with the call to action text/button for picking a primary care provider (PCP). Our direct success metric can be the click through rate of this button, which can roll up into the percentage of members who see a high performing PCP that we recommend, ultimately leading to reduced total cost of care.
Don’t just devise success metrics: establish failure thresholds for these metrics as well. Building on the example above, a success metric could be +5% relative increase for the call to action text change, and failure would be considered a ≤ 2% increase.
Articulate failure criteria for unexpected behavior from an experimental change. Doing this will get your teams explicitly thinking about possible side effects of your experiment. In the example of experimenting with a call to action for selecting a PCP, we wouldn’t want to see an increase in inbound calls to our Care Guide teams as a result of this change.
Ensure that your control & experiment populations have proportional representation of key user segments. As an example, Jaime strives to include proportional representation of Spanish speaking members in control groups when doing product marketing email experiments.
When experimenting with processes and programs, you may not have such clear cut metrics, and may need to resort to proxy metrics or working backwards to devise a way to quantify otherwise qualitative feedback about a process change.
If experimenting with an internal tool, you may not have enough volume to reach statistically significant values in a reasonable timeframe. Both Alan and Brian raised that there’s a balance to statistical rigor versus the realities of testing in an actual business production environment. You won’t always be able to perfect an experiment design, and that’s OK, so long as you’re able to obtain useful data to help inform decisions about your experiment.
Case study: Oscar mobile app home screen experiment
Herry and team recently completed a successful experiment of introducing action items to the mobile app home screen. Action items prompt things that members should do to keep on top of their care and make the most of their coverage, things such as choosing a Primary Care Provider, updating their health profile, reviewing a summary of their recent care, or paying their monthly premiums on time.
Side-by-side comparison of the experiment V.S. control mobile home screens
The primary metric monitored for this experiment was the click through rate on the action items themselves. Herry and Alan worked to establish a success threshold being 3% or greater click through rate, and defined the failure threshold as being less than 2% clickthrough rate.
As it turned out, we saw this click through rate go as high as 50%+, and as a result, the decision was made to eventually ramp the experimental changes up to 100%. A number of secondary metrics were also monitored, primarily in consideration of not creating downside impacts on other calls to action that are in the mobile homepage, and with the goal of avoiding inbound calls to our Care Guides as a result of these changes.
Closing thoughts
At Oscar, experimentation is a powerful tool that enables us to “Seek Insight” into how people interact with various experiences we create. Experimentation lets us verify our hypotheses about a change, whether it’s to an internal product, or a user-facing one, or even a team’s processes, and lets us gauge the true impact of our work. We hope this exploration of experimentation at Oscar inspires you to think creatively about how you might apply experimentation to your own context and practice!