Evidence-Based Coaching Frameworks in Practice

The most rigorous recent data comes from a 2023 meta-analysis published in the Academy of Management Learning and Education. De Haan and Nilsson examined 39 randomized controlled trial samples totaling 2,528 participants and found a statistically significant effect size of g = 0.59 across leadership and personal outcomes. Moderate. Real, but not the revolution the sales decks suggest. The authors also flagged what the field quietly buries: publication bias distorts this literature in a specific direction. Randomized controlled trials are logistically and ethically difficult to run in workplace settings, so the studies that exist skew toward favorable conditions, favorable populations, and results worth publishing. The true population effect is smaller than the headline number implies, and anyone presenting the headline number without that caveat is telling you something about their incentives.
A separate 2023 meta-analysis in Frontiers in Psychology reinforced one finding worth isolating: coaching is especially effective for behavioral change, and that benefit extends well beyond the executive suite to production-level employees. Organizations that scope coaching exclusively for senior leaders are leaving the most durable part of the evidence untouched.
Self-reported data, while methodologically weaker, runs in the same direction. The 2024 ICF/PwC survey across 64 countries found 87% of clients reporting positive return on investment, 80% improved self-confidence, and 70% improved work performance. Perception data, not controlled measurement. Still worth treating as signal.
Two figures circulate constantly in this field and both deserve scrutiny. The 788% ROI number comes from a single Fortune 500 company in a MetrixGlobal study. It cannot generalize. Citing it as a benchmark is a tell: whoever is quoting it is selling something. The more defensible planning figure is the broader 7x average from ICF research, drawn across a wider population. Even that reflects what coaching produces under competent implementation, not what it reliably delivers by default.
The evidence supports coaching as effective. The effect sizes are moderate and sensitive to context. Think of framework selection and application fidelity as the soil in which research findings either take root or wither — they are not procedural details but the conditions that determine whether outcomes materialize at all.
The core frameworks practitioners actually use and what differentiates them
GROW is the oldest and most studied single coaching model in practice. Developed by Whitmore and colleagues in the mid-1980s and published formally in 1992 in "Coaching for Performance," it structures sessions around Goal, Reality, Options, and Will. Its research base is the largest of any individual framework. When you are trying to separate evidence from enthusiasm, that matters more than practitioners usually admit. I have watched coaches dismiss GROW as basic, reaching instead for newer frameworks with better branding, and then produce work that would have benefited from the discipline GROW imposes. Familiarity is not the same as mastery — and knowing the map is not the same as knowing the terrain.
OSKAR, developed by McKergow and Jackson, represents a genuine philosophical departure. It is solution-focused: attention stays on what is already working and on what success looks like, rather than on diagnosing the problem. That orientation is not cosmetic. It changes which questions get asked, where the coachee's attention lands, and what the coach treats as relevant data. OSKAR favors brevity and forward momentum, which is a real strength in certain contexts and a real liability in others.
CLEAR, developed by Peter Hawkins in the early 1980s, was designed for relational depth and team-level work. The stages, Contracting, Listening, Exploring, Action, and Review, unfold more slowly and more expansively than GROW's four-stage structure. That pace is intentional, calibrated for work that requires genuine exploration rather than structured problem-solving. Coaches who find CLEAR uncomfortable are often coaches who find sitting with ambiguity uncomfortable, which tells you something important.
Beyond these three, there are named frameworks with thinner independent research trails: TGROW, ARROW, SCORE, RISE, ACHIEVE, STEPPA. Practitioners who work primarily from frameworks that exist in marketing materials rather than peer-reviewed literature warrant scrutiny. The absence of independent validation is diagnostic information, not an administrative gap.
Why no single framework works across every context, and how to match the right one to the situation
GROW fits structured, goal-focused sessions. It is the right instrument for shorter-term performance coaching engagements, roughly three to six months, where the coachee arrives with a defined challenge and wants a clear path through it. Its directional momentum is a feature when the work is that kind of work. It becomes a constraint when the work is not.
OSKAR's strength is speed and agency. It is well suited to solution-focused brief interventions where dwelling on root causes would actually impede progress, where the coachee already understands what went wrong and needs traction toward what comes next. Its real limitation: it can sidestep confrontation that is genuinely necessary. A coachee who needs to examine a behavioral pattern that keeps recreating the same problem will find OSKAR's forward orientation prematurely comfortable. Moving on is not the same as resolving anything, and a skilled coach knows the difference, even when the coachee prefers the former.
CLEAR's depth makes it the better choice for longer developmental engagements and team-level work. Its pace frustrates action-oriented leaders. That friction is not a flaw in the model; it is a mismatch. Coaches who confuse the two either abandon a useful framework prematurely or apply it to the wrong context and blame the method.
The matching decision turns on at least three variables: the coachee's natural orientation toward problem analysis versus solution-building, the time horizon of the engagement, and the organizational context, whether this is executive coaching, frontline development, or team-level work. ICF's core competency framework makes the principle explicit: coaching presence and flexibility take precedence over fidelity to any single model. The right framework is the one that serves the coachee's actual need in this moment, not the one the coach reaches for out of habit.
How experienced coaches blend frameworks across an engagement rather than applying one end to end
Integration in practice looks like this: using CLEAR's contracting structure to open an engagement, running individual sessions through GROW's four-stage format, applying OSKAR's scaling questions to calibrate progress between sessions, and drawing on ACHIEVE's evaluation criteria when complex decisions surface. The blend is not arbitrary eclecticism. Each framework's distinct mechanism serves a different functional need as the engagement evolves.
A common pattern, one I have seen repeatedly across different organizational contexts: the engagement opens goal-focused, GROW is the right instrument. Mid-engagement, unexpected developmental material surfaces, old assumptions, identity-level questions, relational dynamics the coachee did not anticipate and the coach did not expect. CLEAR's exploratory listening becomes more useful than GROW's structured progression. Later, as the work moves toward consolidation, OSKAR's scaling questions provide a low-friction way to track how far the coachee has actually moved and where the remaining gap sits. The sequence is not planned at the outset. It is recognized as the work unfolds, which is the part no certification program teaches you directly.
Framework knowledge, in this context, is a prerequisite, not a competency. The actual skill is understanding the theoretical mechanism behind each model well enough to know which lens serves the present moment. Knowing the acronym is table stakes. Knowing why OSKAR's forward orientation increases perceived agency while GROW's reality stage can temporarily reduce it, and deploying that knowledge in real time, is what separates experienced practitioners from credentialed novices.
For organizations buying coaching services, "which model does this coach use?" is less diagnostic than asking how the coach decides when to shift approaches. The second question tells you significantly more about actual capability.
The theoretical foundations that explain why these frameworks produce results when they work
Self-Determination Theory, developed by Deci and Ryan in 1985, is the most directly applicable explanatory framework in the literature. Coaching's effectiveness is meaningfully linked to how well the engagement satisfies three basic psychological needs: autonomy, competence, and relatedness. More autonomy-supportive coaching styles are consistently associated with greater need satisfaction, stronger intrinsic motivation, reduced intentions to disengage, and better performance outcomes. This explains why both CLEAR's relational depth and OSKAR's agency-focused questioning produce results: different mechanisms, same underlying need satisfaction.
Positive psychology adds a complementary layer. Coaching's structural emphasis on strengths and future states rather than deficit correction aligns with what the behavioral change literature shows about what sustains new behavior over time. Deficit-focused work activates threat responses. Future-focused work activates approach motivation. There are neurological consequences to that distinction, and they surface in session outcomes in ways that are not subtle once you know what you are looking for.
The theoretical grounding also illuminates a specific and common failure mode. A directive coach who gives correct advice will produce short-term clarity while undermining the coachee's sense of autonomy. The advice is right; the mechanism is wrong. The session feels productive. The behavioral change does not materialize. Frameworks applied without attention to psychological need satisfaction become procedural checklists that produce compliance. Compliance and genuine change are not the same thing — compliance is following the map, genuine change is knowing the territory — and organizations that cannot distinguish between them have built their measurement systems around the wrong variable.
How coaching frameworks translate to measurable outcomes in organizational settings
The 2023 ICF HCI report found that 72% of respondents acknowledged a strong correlation between coaching and increased employee engagement. Engagement is a metric organizations track consistently, budget against, and understand in operational terms. That correlation translates coaching into a language finance and HR leadership can both use without translation.
Research from Dion Leadership in 2024 identified something more specific and more commercially significant: leadership coaching's biggest single-year mover was the desire to stay with the organization, up 7 percentage points from the prior year. Retention is where the cost conversation becomes concrete. The expense of replacing a mid-level or senior employee is a well-documented figure in most HR functions. Coaching that demonstrably moves retention intention belongs in the retention budget, not the development budget. The categorization shapes how it gets evaluated and defended, which shapes whether it survives budget cycles.
Intel's ICF Prism Award-winning coaching program, recognized in 2022, is the most prominent published corporate case in the field. Intel's own reporting attributed hundreds of millions of dollars per year in operating margin to the program. That number comes from one organization and cannot be extrapolated as a sector average. But it marks the plausible ceiling of what structured, scaled coaching produces when implementation is serious and measurement is rigorous.
Evidence from healthcare and public-sector contexts adds texture beyond the corporate data. Coaching in those environments reduced workplace tension, increased personal accountability, improved decision-making capacity, and was associated with reduced absenteeism, outcomes that map directly onto the behavioral change finding from the Frontiers in Psychology meta-analysis. The mechanism travels across contexts even when the presenting problem looks entirely different on the surface.
Outcome quality in organizational settings depends heavily on implementation fidelity. The same framework, applied inconsistently by underprepared practitioners without clear contracting and defined success criteria, produces results that look nothing like the headline studies. The research is measuring best-practice conditions. Most organizational deployments do not come close.
What consistent, correct application actually requires from practitioners and the organizations that deploy them
Framework knowledge is necessary and not sufficient. The 2025 ICF Global Coaching Study found that 60% of coaches now offer training alongside coaching services and 57% offer consulting. Practitioners who understand only one methodology are already behind what the market expects and what clients increasingly need. Versatility is not a differentiator at this point; it is an entry condition for serious organizational work.
Application fidelity requires four things practiced consistently: clear contracting before each engagement, conscious framework selection rather than habitual default, active monitoring of progress against the chosen model's own evaluation criteria, and willingness to shift the framework when it stops serving the coachee. The list reads simply. None of it is simple. The contracting step alone, done properly, requires the coach to surface assumptions the coachee does not yet know they have, which is harder than it sounds and rarer than it should be.
There is a direct analogue in organizational measurement to the publication-bias problem in the research. Coaching programs that measure only satisfaction scores are doing the equivalent of publishing only positive results, capturing what is easy to collect rather than what the evidence says matters: behavioral change, goal attainment, movement on retention metrics. The organizations that get closest to the outcomes the research promises are the ones measuring what the research actually measured.
The field has grown substantially in the number of practitioners since 2019. More choice and more quality variability arrived simultaneously. For organizations making procurement decisions, the question is no longer whether coaching works. It is which practitioners, matched to which frameworks, deployed with which populations, measured against which outcomes. Asking for evidence of framework training, independent validation of the approach, and a clearly articulated outcome measurement methodology is how the research literature becomes a decision that actually produces what the research was measuring.


