
You've probably seen this movie. A candidate shows up with a polished résumé, a tidy LinkedIn, maybe even a pristine IELTS score, and then sends a follow-up message that sounds like it was assembled by a tired vending machine. That's the gap. Not intelligence. Not effort. English proficiency testing sits right in the middle of that gap, and if you hire across borders, it's one of the few tools that can stop you from confusing “looks fluent on paper” with “can do the job.”
The mistake many teams make is treating English as a vibe. It isn't. It's a work skill, and work skills need a floor, a benchmark, and a real-world check. If you get that wrong, you don't just miss a few nice candidates. You ship bad customer emails, waste manager time, and make everyone sit through awkward meetings where nobody wants to admit the handoff broke because the language screen was basically decorative.
A recruiter once told me a candidate was “perfect, except for the writing.” That's not a small exception. That's the part where your customer support lead can't explain a refund, your engineer can't push back in a code review, and your sales rep turns a simple objection into a twenty-minute ambiguity festival. Remote hiring magnifies that problem because you often don't see the bad fit until after onboarding, when the polite messages start hiding the underlying issue.
English proficiency testing exists to catch that before the offer letter. It's not there to judge personality, charisma, or technical talent. It's there to answer a narrower question, can this person read, write, listen, and speak well enough to handle the communication load of the job without creating avoidable friction?
That distinction matters more than people think. A candidate can be brilliant and still struggle with live back-and-forth, fast written context switching, or role-specific vocabulary. A generic “good English” impression won't save you when the client asks for clarification and your new hire answers like they're still in exam mode.
Practical rule: If the role has external customers, cross-functional dependencies, or async collaboration, treat language testing like a working-condition check, not a box to tick.
The historical scale of modern benchmarking is also why this matters. The EF English Proficiency Index has tracked adult non-native English ability since 2011, and its 2025 edition was calculated from test data from 2.2 million test takers in 2024 across 123 countries and regions. It ranked the Netherlands #1 with a score of 624, and Copenhagen led cities at 644 (EF EPI overview). That kind of dataset exists because the market keeps needing a shared reference point, even if the hiring use case is smaller and messier.
The punchline is simple. If you skip proficiency testing, you're not being flexible, you're gambling. And the bill usually shows up later, in customer embarrassment, manager rework, and the kind of attrition nobody wants to explain in a retrospective.
English tests are not one thing. They measure a bundle of abilities, and the bundle matters more than the badge on the cover page. At minimum, a real proficiency test looks at reading, writing, listening, and speaking, then layers in grammar, vocabulary, and pragmatic judgment, the part that decides whether someone sounds clear in a meeting or vaguely alarming in a Slack thread.
A good way to think about it is this. Placement tests sort learners into levels. Aptitude tests try to predict how easily someone might learn. Proficiency tests measure what someone can already do. Recruiters mix these up all the time, which is how they end up using a university-admissions tool for a customer success role and acting surprised when the results don't map cleanly to the job.
The more modern systems also care about how scores are built. In the ACCESS for ELLs technical report, ability, item difficulty, and step parameters are estimated in logits and then converted to reported scores, and item duration can be used to compute item efficiency as information divided by average time (Michigan ACCESS technical report). That's the boring technical bit that matters, because it separates measurement precision from simple test speed.

A useful test gives you a score you can interpret without a forensic accounting degree.
For hiring, that means you want more than a raw total. You want a stable scale, a clear band, and ideally a label that maps to something your team can reason about, like CEFR or ILR. If the report can't be translated into a shared hiring language, then the test is just expensive noise.
For a practical hiring lens, LatHire's guide on how to assess communication skills is useful because it pushes teams toward role-relevant judgment instead of pretending every role needs the same English profile. That's the right instinct. The best tests don't just measure English, they measure whether the candidate can use English in the kind of work you pay them to do.
A recruiter gets a clean score report and still makes a bad hire. That happens when the test measures English in one setting and the job needs it in another. TOEFL iBT, IELTS Academic, Duolingo English Test, and Cambridge C1 Advanced all measure English, but they stress different parts of it. Use them as if they are interchangeable, and you'll optimize for the wrong behavior while calling it rigor.
TOEFL iBT is strongest for academic reading, listening, and integrated writing. IELTS Academic is widely used for admissions and immigration pathways, and its speaking format still reflects a formal test interaction. Duolingo English Test works well for high-volume screening because it is online and fast to deploy. Cambridge C1 Advanced is a serious credential for advanced general English, but it is still an exam, not a simulation of your team's daily workflow.
Workplace-focused assessments solve a different problem. Versant, PTE, and bespoke AI-driven simulations are built more for operational use, especially when you care about spoken interaction, speed, and repeatability. Use them when the communication risk is practical, not academic. A customer support rep who cannot handle a refund conversation needs a different assessment design than a researcher who needs to summarize a paper.
The selection question is simple. Match the test to the failure you want to prevent. That is the frame how to select assessment tools should push you toward, because the decision is always about fit.
Here is the blunt version.
If the role is conversational, do not rely on a test that mostly rewards reading and test familiarity. If the role is documentation-heavy, do not pay for speaking theatrics that never show up in the job.
LatHire's own guidance on English requirements for remote hiring points toward role-relevant tests, short async writing tasks, and live conversation with constraints. That is the direction smart teams are already moving in. The point is not to crown one format. The point is to match the assessment to the communication risk.
CEFR and ILR are two different scoring systems, but they answer the same hiring question, what can this person do in English? CEFR runs from A1 to C2, while ILR runs from 0 to 5. If you hire across borders, you need a shared reference point because a raw test score means very little on its own. A recruiter cannot make a solid call from an isolated number unless that number maps to a level the hiring team understands.
The practical move is to translate scores into bands your team can use in decisions. The University of California publishes specific English proficiency paths for international applicants, including IELTS Academic 6.5+, TOEFL iBT 80+ before January 2026 or 4.5+ starting January 2026, Duolingo 115+, ACT English Language Arts 24+, SAT Writing and Language 31+, or a grade of C or better in a UC-transferable English composition course (UC English language proficiency requirements). That is not a single-test standard. It is a set of acceptable thresholds.
| CEFR Band | Plain English | IELTS | TOEFL iBT | Duolingo | Cambridge |
|---|---|---|---|---|---|
| B1 | Functional, simple workplace English | Lower bands vary by institution | Lower bands vary by institution | Lower bands vary by institution | Lower bands vary by institution |
| B2 | Solid independent use, handles most work communication | Around 6.0 to 6.5 depending on policy | Around 72 to 80 depending on policy | Around 105 to 115 depending on policy | Around 173 to 176 depending on policy |
| C1 | Strong professional fluency, handles nuance | 7.0 and above in many policies | 95 and above in many policies | 120 and above in many policies | 180 and above in many policies |
| C2 | Very advanced, near-native control in many contexts | Top-end bands | Top-end bands | Top-end bands | Top-end bands |
Treat the table as a translation layer, not as a universal rulebook. Institutions and employers set different cut scores, and those differences matter. Score validity windows matter too. Australia's Department of Home Affairs requires approved test results taken within the 3 years before a visa application, and it evaluates proficiency across four components, not as one blended impression. NYU uses a two-year window for most applicants, with explicit exemptions for applicants who completed their last three years of schooling entirely in English (NYU English language testing).
The operational lesson is plain. Translate the score, check how old it is, and check the subskills behind it. A strong total score with one weak band still signals risk in a hiring context.
A company-wide “everyone needs B2” policy sounds clean until it starts rejecting good people for the wrong job. That's not rigor. That's laziness in a blazer. A sales rep and a backend engineer do not need the same English profile, and pretending otherwise wastes time on both sides of the hiring funnel.
Customer support and SDR roles usually need a strong B2 in speaking and listening, because those people live in live objections, clarifications, and de-escalation. Engineers and back-office roles can often function at B1 to B2 if reading and writing are solid and most collaboration is async-first. Content, marketing, and client-facing roles usually need C1 because the work itself is language-heavy. Leadership and customer-comms roles should sit at C1 to C2 with a high speaking floor, because ambiguity in those jobs gets expensive fast.
That's the rule set. Then test it against behavior.

Rule of thumb: Use proficiency testing as the floor, then layer a role simulation on top. Otherwise you're measuring English in a vacuum, which is a great way to hire somebody who tests well and works oddly.
That's also why a generic cutoff causes pain. A company can accidentally reject articulate but ineffective hires for one role, while letting through technically sharp people who can't communicate in the other. The fix is simple, and yes, it takes a little discipline. Set the floor by role family, then use a task that resembles the actual job.
A lot of testing bias hides behind the word “objective.” It sounds clean, but global English isn't clean. Dominant tests like TOEFL and IELTS can carry native-speaker norms and Western testing culture baked into their design, which means a strong candidate from an Outer- or Expanding-Circle context can get dinged for sounding different, not for being unclear. That's a real problem, not a theoretical one.
The easiest way to see it is to stop treating accent as the proxy for competence. In global teams, people use English to ship work, answer customers, write specs, and negotiate meaning across time zones. That doesn't always look like prestige-accent test prep, and it shouldn't have to. A person can be perfectly understandable and still not match the narrow sound of a standardized exam room.
The research gap here is active, not academic trivia. A 2026 validation study of UB-TEP in Indonesia frames localized testing as a response to the limits of global exams and pushes for culturally responsive alternatives aligned with Global Englishes and CEFR (UB-TEP validation study). That's the right direction because it acknowledges something hiring teams often ignore, language is global, but testing culture often isn't.
There are practical ways to compensate without throwing out measurement.
A smarter policy is also easier to defend in front of skeptical stakeholders. You can say, with a straight face, that you care about whether someone can do the work in your operating context. That's a much better position than pretending one exam format is magically culture-free.
Remote hiring falls apart fast when language checks happen too late. Set the floor before the job post goes live, screen early, and save live interview time for candidates who can do the work. If you wait until the final round to assess English, you have already spent interviewer time on the wrong pipeline.
Write the role band first. If the job needs C1 speaking and B2 writing, make that decision internally and build the process around it. Then place a short screening test near the top of the funnel. Keep it light. The goal is to filter out obvious mismatches before they take interview slots from stronger candidates.
Use a proctored speaking step for finalists only. For roles that are sensitive, regulated, or tied to immigration, use approved delivery and identity checks, because the rules are fixed. The UK Home Office accepts a Secure English Language Test only if it is on the approved list, taken at an approved location, and awarded within the two years before application (UK SELT guidance).
A useful hiring policy starts with role-specific requirements, not a generic language slogan. LatHire's English language requirements page is a good example of how to tie language expectations to the actual job, especially when remote collaboration and cross-border hiring are already part of the workflow.

Use AI-driven assessment carefully. Adaptive testing can cut candidate time. Automated speaking scores can standardize a review. Cheating detection and plagiarism checks help when submissions are unsupervised. The problem is false precision, especially when a hiring team treats a score as if it explains everything.
Operational rule: Put band labels and subskill breakdowns directly into your ATS. If hiring managers have to open a PDF and squint, the process is already too clunky.
LatHire's own pre-employment workflow focuses on skills evaluation and role-relevant screening, which is the right frame for English too. English should sit inside the hiring decision, not float above it as a separate ceremony. Trigger a re-test when the score expires, show the band next to the rest of the scorecard, and treat language as one input among several, not the whole decision.
A hiring team that wants usable English testing should stop treating it like a checkbox. Start with one standardized test for comparison, then add one role-based simulation for the kind of writing, speaking, or reading the job requires. Use CEFR floors by role family, not a single company-wide slogan, and write them down before anyone reviews candidates.
If you want a practical next step, start by rewriting one job description with a CEFR floor and a clear testing plan, then tie it to your pre-employment skills testing workflow. Capture subskill breakdowns, because one total score does not tell you enough to make a hiring decision with confidence.
Use the score window as a hard rule, not a guideline. A stale result is a stale result. Audit for fairness across English varieties, because “objective” scoring still carries accent and culture bias. Put the band label into your ATS so hiring managers see the result next to the rest of the scorecard instead of opening a PDF and guessing.
A short version for the wall.

If you want to tighten this fast, rewrite one role with a CEFR floor, add a short async writing task, and run a live speaking check for finalists. Standardized tests are useful for comparability. Role-based work samples are better when the job depends on communication under real constraints. Treat both as inputs, not verdicts, and stop letting a TOEFL score pretend it knows more about job readiness than your hiring process does.
If you are cleaning up remote hiring this week, update one role, add one work sample, and make sure the proficiency result lands in your ATS instead of a forgotten spreadsheet.
