I spent four years at The Lab at OPM redesigning USAJOBS, and the whole time I was staring at a problem I could not fix from inside a job board. We could guide a job seeker to the right announcement, help them understand it, and get them cleanly through apply. Then their application disappeared into a process we did not control — and often nobody got hired at all. The bottleneck was never the front door. It was the assessment.
As of March 2023, 48% of all competitive examining certificates were thrown out — hiring managers looked at the list of “qualified” applicants HR handed them and chose no one. For technical positions the rate was worse. Hires made through competitive public job announcements had fallen to 20% by FY17. Managers had learned to route around the competitive process entirely, because a process that asks applicants to rate themselves and asks HR specialists to evaluate skills they have never practiced does not reliably produce people who can do the job.
Subject Matter Expert Qualification Assessments (SME-QA) was the fix: put actual practitioners — the engineers, the data scientists, the CX strategists — in the assessment seat alongside HR, using structured resume reviews and interviews instead of self-rated questionnaires. And critically, let HR adjudicate veterans' preference only after the technical assessment, so preference is applied to a list of people who can actually do the work. No new law required. It was legal the whole time.
How I got onto the project
By 2018 I had spent three years at The Lab watching this problem from the front door. I had co-created a journey map of the front- and backstage federal hiring experience — job seekers, applicants, HR specialists, and hiring managers on one page — and taken it to OPM leadership including the Director. I had run ethnographic research with more than 60 hiring managers for the Hiring Manager Advisor prototype, a milestone of CAP Goal 3 (opens in new tab). I knew exactly where the process broke and I could not get at it from USAJOBS.
Buried in that research was the shape of the answer. When we asked hiring managers about the tools available to them, they told us that SME panels were “believed by some to be unlawful,” that panels had worked for the few who tried them, but that most managers either did not know they existed or had struggled to convince their agency of their lawfulness. That is a remarkable finding. The technique worked. It was legal. And the thing standing in its way was not policy but belief. Three years later the SME-QA team would still be fighting that exact sentence.
Getting USDS to pick this up took a campaign. Arianne Miller, the Director of The Lab, and Sean Baker went to Eddie Hartwig, then Deputy Director of USDS, repeatedly, asking him to take up federal hiring. The instrument they used was Improving hiring outcomes, a report I wrote with my Lab colleagues Ben Winter, Han Wang, Jen Kaczor, Min Chung, and Tim Vienckowski — four years of USAJOBS research condensed into the case for why the remaining problems were systemic and could not be solved inside a single program office. It took some doing. An earlier draft is what finally moved him. The version linked above is the last one, published in February 2019 after the USDS sprint had concluded — which is why it closes with The Lab's proposal for how to help scale the pilot that resulted.
The argument against USDS taking it on was that hiring is not citizen-facing work. Our counter was that every citizen interaction is downstream of who the government manages to hire — especially in technology, where government has always struggled to recruit the best people. There was a self-interested version of the argument too, and it was the one that landed hardest: The Lab, USDS, 18F, and every other tech-focused innovation shop in government had been hiring through the same excepted service Schedule A workaround. We all knew the competitive process did not work, because none of us used it. Paving that path so other agencies could do the same seemed like the obvious move.
Eddie assigned it to Stephanie Grosser. Once we started talking she got excited, and it was Stephanie who pushed to go after the competitive service itself rather than widen the Schedule A exception. I agreed with her, and we both knew it was the much harder nut to crack. Stephanie is a force of nature; other people at USDS caught the same enthusiasm for fixing federal hiring, and I ended up bringing what I had learned to a genuinely talented team.
USDS ran its own two-week discovery sprint in November 2018 and concluded that the competitive process can be executed in a way that hands hiring managers a short list of truly qualified candidates. Of its five recommendations, the second was the whole ballgame: subject matter experts should evaluate their peers' technical abilities before preference is applied.
So I started working with the USDS team on this while I was still at The Lab. Through the spring of 2019 — the HHS and DOI pilots — I was working with them daily. My Lab term was due to end that October, so I applied to USDS and transitioned in June 2019 with time to spare. The job title changed in the middle of the pilots; the work did not.
The hypothesis, and how it was tested
The team's hypothesis was precise enough to be wrong: if SMEs complete the assessments before HR determines anyone qualified, fewer unqualified applicants reach the certificate, and hiring managers will make more selections.
Two agencies were chosen in spring 2019 on deliberately hard criteria — they had to have already received a certificate with no usable candidates despite knowing qualified people had applied, have at least five vacancies for the same role, be using public delegated examination rather than a shortcut authority, and be hiring at GS-12 or above. HHS (CTO and CIO offices, up to 10 GS-13 IT Specialist roles) and DOI's National Park Service (up to seven GS-13 System Administrator roles across three locations) were the first two to qualify.
Eight SMEs at each agency — GS-13 to GS-15 staff currently doing or managing the job being filled, none of them the selecting official — defined the competencies, wrote the assessments, and made the qualification calls.
The process we documented and published. Note who owns each phase
— SMEs hold the two assessment steps, and HR does not touch
veterans' preference until phase five.
The applicant funnel
The shape of the funnel is the whole argument. Fewer people called qualified, more people actually hired. The numbers below, and most of the research that follows, come from the initial pilots case study (opens in new tab, PDF) and final report (opens in new tab, PDF) — the two documents that spent sixteen months in clearance.
HHS
- 164 original applicants
- 103 passed resume review
- 54 passed the first interview
- 36 found qualified — 22%
- 7 selections
The baseline it replaced
- Across the five most recent comparable GS-13/14 hiring actions at HHS, HR specialists found an average of 51% of applicants qualified, and hiring managers made an average of 3 selections.
DOI / National Park Service
- 224 original applicants
- 78 passed resume review
- 38 passed the first interview
- 25 found qualified — 11%
- 13 selections
The baseline it replaced
- In DOI's most recent similar hiring action, HR specialists found 95% of applicants qualified and the hiring manager made zero selections from the resulting certificate.
That DOI pair is the clearest statement of the problem I have ever seen in federal hiring: a process that called 95% of applicants qualified produced nobody worth hiring. A process that called 11% qualified filled every vacancy and then some. Hiring managers made their first selections within 11 and 16 days of receiving the certificates — against a government-wide baseline of 37 days, the longest single phase of the federal hiring process. Both agencies then shared their certificates with other offices once their own vacancies were filled.
Veterans' preference was adjudicated only after SMEs finished. Four of the 36 qualified applicants at HHS were veterans; five of 25 at DOI. The team's conclusion was blunt, and it holds up: veterans' preference complaints are a scapegoat for weak assessment strategies. Assess properly first, and preference gets applied to a list of people who can genuinely do the job.
What I worked on
SME-QA ran from the 2018 discovery sprint through the 2019 pilots, a second round in 2020–2021, government-wide shared hiring actions, a detail into OPM in 2022, and a wrap in 2023. I was one of a small team — Stephanie Grosser, Will Slack, Kelvin Luu, Neil Sharma, Amy Paris, Jenn Noinaj, and Katherine Nammacher at USDS, with Roseanna Ciarlante, Dianna Saxman, and Kim Holden at OPM. These were the pieces I owned or led.
Rebuilt the job announcement
The job announcement is where federal hiring either sets expectations or destroys trust. The pilots needed an announcement that read like a job, not like a position description pasted into a form. This piece was mine end to end — I was the sole designer on it, ran the research behind it, and built the template (opens in new tab) in plain language, styled closer to a private sector posting. Four decisions did the work:
- Responsibilities described the duties SMEs defined in the job analysis workshop — not the text of the position description.
- Qualifications listed the technical competencies SMEs had identified as genuinely required.
- How You'll Be Assessed stated plainly that SMEs would review only the first two or three pages of work history. Telling applicants the rule is the difference between a limit and a trap.
- Overview declared that the announcement would close at midnight on the day it hit 100 applications.
HHS received 164 applications and DOI 224 — both announcements closed within two days of opening. This template is now the one used on usajobs.gov for all positions, not just SME-QA ones, which makes it the widest-reaching artifact I have shipped in government.
Designed and built smeqa.usds.gov
Results do not spread themselves. HR specialists across 100+ agencies needed a place to see that this was legal, see how it worked phase by phase, and see the numbers. I co-designed and built smeqa.usds.gov (opens in new tab) — a static site walking a hiring team through every phase (opens in new tab) from job analysis workshop to issuing the certificate, with the legal authority (opens in new tab) for each step and public outcome data (opens in new tab) for every action we ran — including the ones that went badly.
Publishing the numbers was deliberate. The team's first case study and final report sat in White House clearance for one year and four months and were stale by the time they cleared. A site we controlled and could update let us put results in front of agencies while they still mattered. Same pattern I used for the USAJOBS Help Center: small, static, open tools, cheap to host, impossible to break under traffic. It is still up years after the program ended, which was the point.
Contributed to the design research
Research ran continuously through both pilots, with subject matter experts, hiring managers, and job applicants. I conducted a share of those interviews rather than all of them — the work was split across the team — and contributed to the synthesis that turned them into process changes. What follows is what we learned, not a list of things I did alone. It is worth reading in that spirit, because the findings are more interesting than the credit.
A workshop I led, with SMEs from across government mapping what they
actually measure and which tools they actually use. Getting the right
people in a room to define what a job requires is the step everything
downstream depends on — and the one both pilot agencies
under-scoped.
- The blind review experiment. To test whether SMEs were actually necessary — or whether HR specialists could get the same result using SME-written criteria — we contracted two independent HR specialists to review all 388 pilot resumes against the identical competencies the SMEs had used. Their determinations diverged from the SMEs' and from each other's. That is the finding that closes the argument: the criteria are not the hard part, the expertise to apply them is.
- Training length changed inter-rater reliability. This one was a teammate's finding, and it is my favourite of the set. At HHS, 27% of resume reviews needed a third SME to break a tie — a signal that SMEs were applying the standards inconsistently. Extending SME training from two hours to three to allow practice and calibration, and eliminating a confusing “borderline” rating, dropped that to 20% at DOI. One hour of training and one fewer option on a rating scale.
- The page limit did not hurt anyone. When we interviewed SMEs after resume review, they confirmed the two-to-three page limit did not impair their ability to judge qualifications, and applicants who submitted longer resumes were not disqualified at a higher rate. OPM has since updated the Delegated Examining Operations Handbook to permit the practice government-wide.
- The real cost to HR. Interviewing the participating HR specialists let us put a number on their contribution: an average of 118 hours per pilot over three months. HHS's decision to put an HR specialist on every interview for legal defensibility added 145 hours on its own and produced no benefit — DOI reviewed SME transcripts afterward instead, and got the same assurance for a fraction of the cost.
- The job analysis workshop was under-scoped. Talking to participating SMEs told us the workshops felt rushed at both agencies. HHS allotted nine hours and DOI 12; SMEs at both said they needed more time testing and iterating the interview questions, given how much weight those assessments carried. We moved the recommendation to a full 16–20 hours — the single largest time increase we asked agencies to accept, and the one with the most leverage, because every later phase inherits the quality of what comes out of that room.
Attacked the scheduling bottleneck
Two frictions consistently stalled hiring actions: corralling resume reviews out of busy SMEs, and scheduling hundreds of interviews by hand. The team identified a cloud-based scheduling tool before the pilots began, but it had not been approved for government use — so round one was scheduled manually by HR liaisons at HHS and by the hiring managers' administrative staff at DOI, with guidelines capping SMEs at three interviews a day and forbidding back-to-back slots. Those guidelines broke under real conditions almost immediately.
The fix that worked at DOI cost nothing: run the two interviews concurrently. Instead of finishing every first interview before starting any second one, schedulers used a shared cloud spreadsheet that HR updated in real time, and booked an applicant's second interview as soon as they passed the first. That single change is most of the difference between HHS's 3.5 months and DOI's 2 months from announcement to certificate.
I later helped procure a self-service scheduling tool so applicants could book their own slots, and teammates prototyped a resume review and tie-breaking tool that let SMEs check off requirements rather than write prose justifications. With both in place, round two pilots reached 5.5 weeks from announcement to certificate.
The Resume Review Tool, as an SME sees it. Every design decision here
is a finding from the pilots: the yes/no per competency instead of a
prose justification, the structured reason codes so HR can audit
without second-guessing, the recuse button for the conflict-of-interest
problem we hit when SMEs came from multiple offices, and the note that
once an applicant misses one competency you can stop reading.
The automation we decided not to build
The volume problem was real. More than 100 to 200 applications overwhelms a panel of eight to ten SMEs, which is why both pilot announcements were written to close at midnight on the day they hit 100 applicants — and why both closed in two days. The obvious way out was to automate resume screening, and in 2019 we looked hard at it. We talked to Amazon and Google, both of whom had tried training algorithms to assess applicant resumes. Both had abandoned the practice, citing algorithms that reinforced existing bias and correlated negatively with what they were meant to predict. We looked at off-the-shelf testing too, and concluded that putting a long assessment in front of scarce senior candidates would mostly cost us the strongest applicants.
So the recommendation was that the federal government keep resume review a human task for competitive technical roles, and that the engineering effort go to the unglamorous part instead: the tie-break workflow, the structured reason codes, the scheduler. Seven years later I spend most of my time building AI tooling, and I still think that was the right call. The automation that paid off was the automation that removed coordination overhead. The automation we declined was the kind that would have replaced the judgement — which was the only part of the process actually producing the result.
Evangelized the process
I was part of a presentation on SME-QA at the Partnership for Public Service. The team also maintained three standing decks — one for USDS audiences, one for government audiences, one for the public — so that any new talk started from an edit rather than a blank page. It is a small operational habit that saved an enormous amount of time over four years, and I have used it on every team since.
What it produced
Every hiring action was published with its outcomes. As of the last update to the public results page:
Across the program
- 17 completed hiring actions
- 42 agencies
- 6,450 applicants assessed
- 933 qualified by subject matter experts
- 406 selections
- 274 accepted offers
The point
- These are people who are actually in the job. The comparison is not “406 versus some other number” — it is 406 versus certificates that got thrown out and positions that stayed empty.
Time to hire collapsed
- HHS, the first pilot: 3.5 months from announcement to certificate, with OPM auditing every stage.
- DOI, immediately after: 2.5 months, via concurrent interviews and dropping the indecisive “borderline” rating.
- Round two, having swapped the first interview for written and asynchronous assessments and added the Resume Review Tool: 5.5 weeks. EPA and CMS made immediate selections.
- CMS's fall 2021 action ran in under five business weeks.
What a hiring action looks like when the tooling is in place: eight
weeks from workshop to offers.
Quality held
- 100% approval rating on hires made through SME-QA.
- The largest action — the government-wide data scientist hiring action (opens in new tab) — qualified 107 data scientists from 513 applicants and produced 105 selections.
- Two people hired through pilots came back around: one joined USDS, and one became USDS's primary partner at their agency.
It opened the door to people outside government
Federal hiring quietly rewards applicants who know how to write a federal resume. Round two tracked employment background through the funnel to see whether SME assessment changed that. At CMS, private sector candidates were 50% of applicants but 67% of those found qualified. Resume length tracking showed the page limit neither advantaged nor disadvantaged either group — though applicants with one-page resumes or resumes over ten pages fared worse than those in between.
There is a postscript to the page limit work. In September 2025 OPM made a two-page limit on resume length (opens in new tab) government-wide policy under the Merit Hiring Plan, and USAJOBS now enforces it at upload. I am not going to claim we caused that — six years and two administrations sit in between. But it is a development worth noting, and our data has something to say about it: we found the sweet spot was two to nine pages, and that applicants with one-page resumes fared worse than the middle of the range. A limit on what reviewers must read is not quite the same instrument as a hard cap on what applicants may submit.
Who applied at CMS
Who was found qualified at CMS
What outlasted us
The most durable outcome was not the process. It was the plumbing. OPM rebuilt our resume review and tie-breaking prototype inside USA Staffing, which means every federal hiring team now has the ability to route resume review and tie-breaking to subject matter experts rather than to HR. That single integration reaches further than any pilot we ran.
OPM also updated the Delegated Examining Operations Handbook to let agencies limit the work history pages reviewed during qualification — a small change that came directly out of what we measured. We made hiring data public, published guidance agencies still use, and stood up a hiring experience office inside OPM (HX, originally scoped as a Hiring Assessment Line of Business) so the work had somewhere to live after USDS left. The qualification definitions from the government-wide CX and data science actions — what it actually means to be a CX strategist, what it actually means to be a data scientist — have been reused by agencies over and over.
What I would do differently
SME-QA did not become the default way government hires, and I think the reasons are instructive.
- Changing handbooks does not change behavior. The team did a full roadshow to agencies and taught a class. Both cost enormous effort. Agencies hear recommendations constantly; we were one more voice. What actually moved an agency was being in the room running a real hiring action with them. OPM made changes to the handbook but never really endorsed SME-QA as a way to hire tech talent. The USDS team didn't scale to be able to run the process repeatedly across agencies and for multiple roles. In retrospect, it may have been better to make our first 2 pilot agencies more successful across roles and then use those as case studies to get other agencies to adopt the process.
- A process that needs a dedicated project manager is a process with a design flaw. Our own case study concluded that SME-QA needs a project manager outside HR to hold the schedule together. That was an honest finding, but in hindsight it was also a warning: we had built something that only ran when we were personally pushing it. The adoption data bears that out.
- Building up credibility in a byzantine system is difficult. USDS had no inherent authority in hiring policy. The team put in57,000 words of policy call notes over two and a half years just to be credible enough in the room to answer “isn't this illegal?” That was the price of admission, not overhead.
- Start clearance immediately. Sixteen months of review made our first case study a historical document. Anything that has to clear should enter the pipeline the day it is drafted well enough to read.
Why this one stays with me
This is the piece of work I point to when someone asks what civic tech is for. Nobody had to change a law. The authority to do this existed the entire time; what was missing was a team willing to learn the policy well enough to prove it, design a process people could actually follow, measure it honestly enough to publish the parts that did not work, and stay with it for four years. The site is still up. The numbers are still public. The tooling is still in USA Staffing.