ABSTRACT
Pseudomonas aeruginosa is a clinically important Gram-negative pathogen responsible for a wide variety of serious nosocomial and community-acquired infections. Antibiotic resistance is a major concern, as this organism has a wide variety of resistance mechanisms, including chromosomal class C (blaPDC) and D (blaOXA-50 family) β-lactamases, efflux pumps, porin channels, and the ability to readily acquire additional β-lactamases. Surveillance studies can reveal the diversity and distribution of β-lactamase alleles but are difficult and expensive to conduct. Herein, we apply a novel approach, using publicly available data derived from whole genome sequences, to explore the diversity and distribution of β-lactamase alleles across 30,452 P. aeruginosa isolates. The most common alleles were blaPDC-3, blaPDC-5, blaPDC-8, blaOXA-488, blaOXA-50, and blaOXA-486. Interestingly, only 43.6% of assigned blaPDC alleles were encountered, and the 10 most common blaPDC and intrinsic blaOXA alleles represent approximately 75% of their respective total alleles, while many other assigned alleles were extremely uncommon. As anticipated, differences were observed over time and geography. Surprisingly, more distinct unassigned alleles were encountered than distinct assigned alleles. Understanding the diversity and distribution of β-lactamase alleles helps to prioritize variants for further research, select targets for drug development, and may aid in selecting therapies for a given infection.
KEYWORDS: beta-lactamases, antibiotic resistance, Pseudomonas aeruginosa, bioinformatics
INTRODUCTION
Pseudomonas aeruginosa is a clinically important, Gram-negative pathogen responsible for a wide variety of both nosocomial and community-acquired infections, including burn, wound, respiratory tract, urinary tract, and bloodstream infections (1) and commonly harbors high levels of antibiotic resistance (2). Carbapenem-resistant P. aeruginosa (CRPa) is considered a “high-priority pathogen” by the World Health Organization (WHO) (3), and the Centers for Disease Control and Prevention (CDC) has declared multidrug-resistant (MDR) P. aeruginosa a “serious threat,” responsible for 2,700 deaths and $767 million in healthcare costs in the United States in 2017 (4). Likely driven by the COVID-19 pandemic, rates of hospital-onset MDR P. aeruginosa infections increased 32% from 2019 to 2020 (5).
Unfortunately, treatment of these diverse infections is complicated by a wide array of intrinsic, acquired, and mutational mechanisms—including multiple β-lactamases, alterations in outer membrane porins and efflux pumps, and target modification—that provide resistance to large numbers of antibiotics (2).
While many recent analyses have examined the prevalence of β-lactamases in P. aeruginosa (6–12), such studies often utilize a relatively limited sample size (both in quantity and distribution), have a specific and somewhat limited focus (e.g., carbapenemase genes), look only at data collected two or more years prior to publication, and rarely utilize whole-genome sequencing (WGS) as a means of exploring the full diversity of the resistome. A database of amino acid variants associated with assigned blaPDC alleles has been developed (13), along with a P. aeruginosa diversity panel, which captures a wide variety of β-lactamase alleles and other resistance determinants (14), but neither provides dynamic snapshots of allelic diversity.
Given the relative lack of broad surveys of β-lactamase alleles, we set out to gain a better understanding of the genetic landscape of P. aeruginosa. Herein, we analyzed data on more than 30,000 isolates found in the National Center for Biotechnology Information (NCBI) Pathogen Detection databases to gain insights into the P. aeruginosa β-lactamase resistome. Querying and analyzing a large, rich data set allow for the examination of a wide array of variables and provide additional dimensions over which to analyze the data. Furthermore, the NCBI Pathogen Detection pipeline provides data in near real time (15), potentially allowing changes in patterns of resistance genes to be examined quickly and efficiently. Understanding the genetic diversity of β-lactamases in P. aeruginosa will help drive innovation in multiple ways: at the basic science level, it will help prioritize β-lactamase variants to study biochemically and microbiologically to understand changing resistance mechanisms; at the translational level, it will help select targets for the design, optimization, and testing of new β-lactam antibiotics and β-lactamase inhibitors; and at the clinical level, it can aid in the selection of the most promising therapies for a given infection.
RESULTS
Overview of the data set
As of 9 September 2024, the NCBI Microbial Browser for Identification of Genetic and Genomic Elements (MicroBIGG-E) database covered 31,061 P. aeruginosa isolates. After filtering for isolates encoding at least one full-length, high-quality blaPDC allele and one full-length, high-quality intrinsic blaOXA-50 family allele, and deduplicating at the isolate level, the final data set contained 30,452 isolates encoding a total of 73,147 distinct β-lactamase (bla) genes (Table S1).
Unsurprisingly, blaPDC (30,842 genes) and intrinsic blaOXA (blaOXA-50 family members; 30,571 genes) were the most commonly encountered bla gene families, collectively accounting for 84.0% of all bla alleles in the data set. Members of 9 acquired blaOXA families and 28 other acquired bla gene families were encountered in five or more isolates each, representing a total of 10,792 acquired genes. For the purposes of this analysis, “assigned” refers to a distinct bla allele that has been designated with an identifying number (e.g., blaPDC-3); “unassigned” bla alleles do not have a specific number attributed to that bla gene. Collectively, 531 distinct, assigned bla alleles were present, including 254 distinct, assigned blaPDC alleles and 63 distinct, assigned, intrinsic blaOXA alleles. A total of 3,160 unassigned bla alleles were present, including 1,473 unassigned blaPDC genes (622 distinct alleles) and 1,369 unassigned, intrinsic blaOXA genes (262 distinct alleles) (Table S2).
Origin of isolates
Geographic
The data set contains one or more isolates from 99 countries and at least 400 from each of the eight regions examined (Fig. 1; Fig. S1; Table S3). The United States and China were the largest contributors (30.2% and 9.4% of isolates, respectively), with 39 countries contributing 50 or more isolates each (Table S3). North America, Europe, Central Asia, and East & Southeast Asia contributed the most isolates (Table S3; Fig. S1). Unsurprisingly, higher- and middle-income countries tend to be better represented.
Fig 1.
Geographic origin of the isolates used in the analysis. Geographic origin was unavailable for 4,774 of 30,452 isolates. Gray indicates that no isolates were analyzed from a given country/region. The map was generated using the ggplot2 package for R with the default world map data.
Collection date
Both historic and contemporary isolates are included, but the vast majority have been collected in recent years as large-scale collection, and sequencing has become more frequent and affordable. A total of 495 isolates were collected before the year 2000, and an average of 1,916 isolates were collected annually from 2015 to 2023 (Table S4).
Overview of β-lactamase genes
The three most commonly assigned blaPDC alleles are blaPDC-3 (5,670 isolates, 18.6%), blaPDC-5 (4,269 isolates, 14.0%), and blaPDC-8 (3,577 isolates, 11.7%). An additional 12 alleles occur in 1.0%–10.0% of isolates and 9 in 0.5%–1.0% of isolates (Table 1). Contrastingly, 171 alleles occur in 10 or fewer isolates, including 70 found in only a single isolate (Table S5). Unassigned alleles follow similar patterns, with just one distinct allele (of 622 total) occurring in 0.10% or more of isolates (NCBI Protein Accession Number HDQ8152859.1 [281 isolates, 0.92%]) and 614 distinct alleles occurring in five or fewer isolates, including 513 occurring only once (Table S6).
TABLE 1.
Intrinsic blaPDC (left) and intrinsic blaOXA (right) alleles occurring in 100 or more isolatesa
| Allele | Isolates | Percentage | Rank | Allele | Isolates | Percentage | Rank |
|---|---|---|---|---|---|---|---|
| bla PDC-3 | 5,670 | 18.62 | 1 | bla OXA-488 | 4,672 | 15.34 | 1 |
| bla PDC-5 | 4,269 | 14.02 | 2 | bla OXA-50 | 3,821 | 12.55 | 2 |
| bla PDC-8 | 3,577 | 11.75 | 3 | bla OXA-486 | 3,608 | 11.85 | 3 |
| bla PDC-1 | 2,154 | 7.07 | 4 | bla OXA-494 | 3,126 | 10.27 | 4 |
| bla PDC-35 | 2,137 | 7.02 | 5 | bla OXA-395 | 2,543 | 8.35 | 5 |
| bla PDC-19a | 1,539 | 5.05 | 6 | bla OXA-396 | 2,161 | 7.10 | 6 |
| bla PDC-16 | 1,145 | 3.76 | 7 | bla OXA-847 | 1,524 | 5.00 | 7 |
| bla PDC-34 | 906 | 2.98 | 8 | bla OXA-904 | 845 | 2.77 | 8 |
| bla PDC-24 | 850 | 2.79 | 9 | bla OXA-905 | 768 | 2.52 | 9 |
| bla PDC-11 | 656 | 2.15 | 10 | bla OXA-846 | 724 | 2.38 | 10 |
| bla PDC-15 | 464 | 1.52 | 11 | bla OXA-848 | 669 | 2.20 | 11 |
| bla PDC-31 | 396 | 1.30 | 12 | bla OXA-851 | 588 | 1.93 | 12 |
| bla PDC-12 | 361 | 1.19 | 13 | bla OXA-1028 | 552 | 1.81 | 13 |
| bla PDC-30 | 331 | 1.09 | 14 | bla OXA-1032 | 485 | 1.59 | 14 |
| bla PDC-36 | 330 | 1.08 | 15 | bla OXA-1035 | 452 | 1.48 | 15 |
| bla PDC-37 | 248 | 0.81 | 16 | bla OXA-914 | 420 | 1.38 | 16 |
| bla PDC-14 | 240 | 0.79 | 17 | bla OXA-1127 | 234 | 0.77 | 17 |
| bla PDC-103 | 215 | 0.71 | 18 | bla OXA-1034 | 212 | 0.70 | 18 |
| bla PDC-6 | 169 | 0.55 | 19 | bla OXA-1026 | 154 | 0.51 | 19 |
| bla PDC-121 | 162 | 0.53 | 20 | bla OXA-901 | 153 | 0.50 | 20 |
| bla PDC-39 | 161 | 0.53 | 21 | bla OXA-1014 | 131 | 0.43 | 21 |
| bla PDC-60 | 161 | 0.53 | 21 | bla OXA-1030 | 126 | 0.41 | 22 |
| bla PDC-22 | 158 | 0.52 | 22 | bla OXA-1033 | 119 | 0.39 | 23 |
| bla PDC-46 | 153 | 0.50 | 23 | bla OXA-906 | 118 | 0.39 | 24 |
| bla PDC-59 | 123 | 0.40 | 24 | ||||
| bla PDC-120 | 120 | 0.39 | 25 | ||||
| bla PDC-66 | 118 | 0.39 | 26 | ||||
| bla PDC-71 | 115 | 0.38 | 27 | ||||
| bla PDC-98 | 110 | 0.36 | 28 |
Percentage represents the percentage of isolates encoding a specific allele.
Among intrinsic blaOXA alleles, the four most common are blaOXA-488 (4,672 isolates, 15.3%), blaOXA-50 (3,821 isolates, 12.5%), blaOXA-486 (3,608 isolates, 11.8%), and blaOXA-494 (3,126 isolates, 10.3%). An additional 12 alleles occur in 1.0%–10.0% and 4 in 0.5%–1.0% of isolates (Table 1). In contrast to blaPDC, only 45 alleles occur in 10 or fewer isolates and just 6 in only a single isolate (Table S5). Three distinct unassigned alleles occur in 0.10% or more of isolates (RefSeq [16] protein accession number WP_058176185.1 [49 isolates, 0.16%], WP_034047652.1 [31 isolates, 0.10%], and WP_034040414.1 [30 isolates, 0.10%]), while 349 appear in five or fewer isolates, including 165 in only a single isolate (Table S6).
Acquired bla alleles represent a total of 10,494 genes, consisting of 174 distinct, assigned alleles across 10 blaOXA families and 27 other bla gene families found in five or more isolates. The most common acquired gene was blaVIM (2,454 isolates), and an additional 12 were found in 100 or more isolates each: blaOXA-10 family, blaOXA-2 family, blaGES, blaIMP, blaKPC, blaNDM, blaCARB, blaOXA-1 family, blaTEM, blaVEB, blaPER, and blaPAC (Table S7). The seven most commonly acquired alleles, each occurring in 1.0% or more of isolates, were blaVIM-2 (1,703 alleles, 5.6%), blaOXA-2 (1,240 isolates, 4.1%), blaOXA-10 (955 isolates, 3.1%), blaKPC-2 (822 isolates, 2.7%), blaNDM-1 (714 isolates, 2.3%), blaCARB-2 (479 isolates, 1.6%), and blaGES-5 (324 isolates, 1.1%) (Table 2). Collectively, members of three blaOXA families appeared in 1.0% or more of isolates: the blaOXA-10 family (1,493 isolates, 4.9%), the blaOXA-2 family (1,319 isolates, 4.3%), and the blaOXA-1 family (477 isolates, 1.6%) (Table S7).
TABLE 2.
Acquired bla alleles found in 100 or more isolatesa
| Allele | blaOXA family | Isolates | Percentage |
|---|---|---|---|
| bla VIM-2 | 1,703 | 5.59 | |
| bla OXA-2 | bla OXA-2 | 1,240 | 4.07 |
| bla OXA-10 | bla OXA-10 | 955 | 3.14 |
| bla KPC-2 | 822 | 2.70 | |
| bla NDM-1 | 714 | 2.34 | |
| bla CARB-2 | 479 | 1.57 | |
| bla GES-5 | 324 | 1.06 | |
| bla IMP-7 | 227 | 0.75 | |
| bla VEB-9 | 224 | 0.74 | |
| bla OXA-4 | bla OXA-1 | 221 | 0.73 |
| bla GES-9 | 221 | 0.73 | |
| bla VIM-1 | 214 | 0.70 | |
| bla PER-1 | 213 | 0.70 | |
| bla VIM-80 | 209 | 0.69 | |
| bla TEM-1 | 205 | 0.67 | |
| bla TEM-116 | 187 | 0.61 | |
| bla IMP-1 | 181 | 0.59 | |
| bla VIM-4 | 163 | 0.54 | |
| bla GES-1 | 162 | 0.53 | |
| bla OXA-246 | bla OXA-10 | 129 | 0.42 |
| bla OXA-1 | bla OXA-1 | 123 | 0.40 |
| bla GES-20 | 121 | 0.40 | |
| bla GES-19 | 117 | 0.38 | |
| bla PAC-1 | 114 | 0.37 | |
| bla OXA-101 | bla OXA-10 | 104 | 0.34 |
| bla IMP-13 | 103 | 0.34 |
Percentage represents the percentage of isolates encoding a specific allele.
Allele frequency by region
Among blaPDC, the most common allele is blaPDC-3 in all regions except East & Southeast Asia (third most common after blaPDC-8 and blaPDC-5) and South Asia (second most common after blaPDC-11). Uniformity across regions is high, with blaPDC-1, blaPDC-3, blaPDC-5, blaPDC-8, blaPDC-19a, and blaPDC-35 generally among the most common alleles in all regions. The most region-specific alleles are blaPDC-16 (second most common in Sub-Saharan Africa but 5th to 13th most common elsewhere), blaPDC-98 (fifth most common in South Asia and no higher than 16th most common in any other region), and blaPDC-97 (seventh most common in Oceania but only 19th to 36th most common in the three other regions it appears) (Fig. 2A; Table S8).
Fig 2.
Intrinsic alleles by region. Rank and proportion of the five most common (A) blaPDC and (B) blaOXA alleles in isolates from each region across all regions. Color indicates the proportion of isolates containing a given allele, and size indicates the frequency rank of an allele in each region.
Among intrinsic blaOXA, blaOXA-488 is the most or second most common in all regions except Sub-Saharan Africa (fourth most common), while blaOXA-486 and blaOXA-395 are among the three most common in six and five of the eight regions, respectively. The most region-specific alleles are blaOXA-914 (9th most common in North America, but 14th to 21st in the six other regions it is found), blaOXA-905 (7th most common in Europe & Central Asia, but 12th to 28th in the other regions it is found), and blaOXA-1029 (9th most common in South Asia, but 14th to 33rd in the other regions it is found) (Fig. 2B; Table S9).
Interestingly, the prevalence of acquired alleles varies widely by region. Acquired blaOXA-10 family genes range in prevalence from 1.9% of isolates in Oceania to 30.7% in South Asia, blaOXA-2 family genes from 0.9% in South Asia to 8.0% in North America, blaVIM genes from 1.5% in Oceania to 21.0% in Sub-Saharan Africa, blaIMP genes from 1.8% in North America to 8.2% in Oceania, blaGES genes from 1.9% in Europe & Central Asia to 12.3% in South Asia, blaKPC genes from 0.1% in Oceania and Europe & Central Asia to 17.5% in South America, and blaNDM genes from 0.9% in Europe & Central Asia to 14.9% in Sub-Saharan Africa (Table S10).
Allele frequency with time
The most common blaPDC allele overall (blaPDC-3) remains the most common, and the second and third overall (blaPDC-5 and blaPDC-8) remain among the top four most common across all periods since 2000, demonstrating the stability of the most common alleles over time. Contrastingly, blaPDC-19a increased from 10th most common in 2000–2012 to 3rd most common in 2022–2024, and blaPDC-24 decreased from 4th most common in 2000–2012 to 9th to 12th most common in all subsequent periods (Fig. 3A; Table S11).
Fig 3.
Intrinsic alleles by collection period. Rank and proportion of the five most common (A) blaPDC and (B) blaOXA alleles in isolates from each collection period across all periods. Color indicates the proportion of isolates containing a given allele, and size indicates the frequency rank of an allele in each period.
Among blaOXA alleles, the six most common alleles overall remain relatively stable over time, ranking among the six most common alleles in all time periods since 2000. Within this, blaOXA-395 increased from sixth most common through 2016 to third most common in 2020–2021 and most common in 2022–2024 (Fig. 3B; Table S12).
Allele combinations
Combinations of β-lactamase alleles are relatively well-distributed, with 21 allele combinations occurring in 1.0% or more of isolates, the five most common of which are blaPDC-5/blaOXA-494 (1,014 isolates, 3.3%), blaPDC-34/blaOXA-488 (822 isolates, 2.7%), blaPDC-5/blaOXA-396 (820 isolates, 2.7%), blaPDC-3/blaOXA-486 (820 isolates, 2.7%), and blaPDC-8/blaOXA-50 (795 isolates, 2.6%) (Table 3). Contrastingly, 1,975 combinations occur in 10 or fewer isolates, including 1,165 in only a single isolate (Table S13). Only three combinations occurring in 1.0% or more of isolates contain an acquired β-lactamase: blaPDC-3/blaOXA-395/blaVIM-2 (13th most common; 528 isolates, 1.7%), blaPDC-8/blaOXA-486/blaKPC-2 (18th most common; 367 isolates, 1.2%), and blaPDC-19a/blaOXA-488/blaNDM-1 (21st most common; 306 isolates, 1.0%) (Table S13), underscoring the relative infrequency of acquired bla alleles in P. aeruginosa. Considering only intrinsic alleles, the most common combinations are blaPDC-35/blaOXA-488 (2,072 isolates, 6.8%), blaPDC-3/blaOXA-395 (1,221 isolates, 4.0%), blaPDC-3/blaOXA-486 (1,185 isolates, 3.9%), blaPDC-5/blaOXA-494/ (1,138 isolates, 3.7%), and blaPDC-5/blaOXA-396 (928 isolates, 3.0%) (Table S14).
TABLE 3.
Allele combinations occurring in 100 or more isolates
| Combination of alleles | Isolates | Percentage |
|---|---|---|
| blaOXA-494, blaPDC-5 | 1,014 | 3.33 |
| blaOXA-488, blaPDC-34 | 822 | 2.70 |
| blaOXA-396, blaPDC-5 | 820 | 2.69 |
| blaOXA-486, blaPDC-3 | 820 | 2.69 |
| blaOXA-50, blaPDC-8 | 795 | 2.61 |
| blaOXA-50, blaPDC-3 | 720 | 2.36 |
| blaOXA-486, blaPDC-24 | 688 | 2.26 |
| blaOXA-847, blaPDC-1 | 634 | 2.08 |
| blaOXA-50, blaPDC-1 | 625 | 2.05 |
| blaOXA-396, blaPDC-8 | 617 | 2.03 |
| blaOXA-905, blaPDC-8 | 606 | 1.99 |
| blaOXA-494, blaPDC-3 | 546 | 1.79 |
| blaOXA-395, blaPDC-3, blaVIM-2 | 528 | 1.73 |
| blaOXA-488, blaPDC-35 | 519 | 1.70 |
| blaOXA-494, blaPDC-15 | 434 | 1.43 |
| blaOXA-848, blaPDC-16 | 434 | 1.43 |
| blaOXA-50, blaPDC-5 | 415 | 1.36 |
| blaKPC-2, blaOXA-486, blaPDC-8 | 367 | 1.21 |
| blaOXA-904, blaPDC-3 | 359 | 1.18 |
| blaOXA-847, blaPDC | 332 | 1.09 |
| blaNDM-1, blaOXA-488, blaPDC-19a | 306 | 1.00 |
| blaGES-5, blaOXA-488, blaPDC-35 | 281 | 0.92 |
| blaOXA-488, blaPDC-19a | 278 | 0.91 |
| blaOXA-395, blaPDC-3 | 259 | 0.85 |
| blaOXA-395, blaPDC-36 | 258 | 0.85 |
| blaOXA-2, blaOXA-914, blaPDC-5 | 247 | 0.81 |
| blaOXA-50, blaPDC-14 | 233 | 0.77 |
| blaOXA-1035, blaPDC-19a | 232 | 0.76 |
| blaOXA-486, blaPDC-5 | 225 | 0.74 |
| blaOXA, blaPDC-5 | 202 | 0.66 |
| blaOXA-488, blaPDC-35, blaVIM-2 | 198 | 0.65 |
| blaOXA, blaPDC-3 | 198 | 0.65 |
| blaGES-9, blaOXA-10, blaOXA-395, blaPDC-19a, blaVIM-80 | 186 | 0.61 |
| blaOXA-486, blaPDC-8 | 185 | 0.61 |
| blaOXA-486, blaPDC-1 | 184 | 0.60 |
| blaOXA-494, blaPDC-8 | 174 | 0.57 |
| blaOXA-851, blaPDC-5 | 174 | 0.57 |
| blaIMP-7, blaOXA-2, blaOXA-846, blaPDC-11 | 167 | 0.55 |
| blaOXA-488, blaPDC-30 | 164 | 0.54 |
| blaOXA-1034, blaPDC-3 | 155 | 0.51 |
| blaOXA-846, blaPDC-11 | 149 | 0.49 |
| blaOXA-395, blaPDC-30 | 147 | 0.48 |
| blaOXA-494, blaPDC | 147 | 0.48 |
| blaNDM-1, blaOXA-395, blaPDC-16 | 146 | 0.48 |
| blaOXA-1028, blaPDC-8 | 142 | 0.47 |
| blaOXA-488, blaPDC-37 | 141 | 0.46 |
| blaOXA-4, blaOXA-486, blaPDC-3, blaVIM-2 | 139 | 0.46 |
| blaOXA-904, blaPDC-5 | 138 | 0.45 |
| blaOXA-1032, blaPDC-16 | 130 | 0.43 |
| blaOXA-2, blaOXA-488, blaPDC-35 | 128 | 0.42 |
| blaOXA-1032, blaPDC-22 | 127 | 0.42 |
| blaOXA-1127, blaPDC-5 | 120 | 0.39 |
| blaOXA-50, blaPDC-31 | 120 | 0.39 |
| blaOXA-1028, blaPDC-5 | 119 | 0.39 |
| blaOXA-851, blaPDC-8 | 115 | 0.38 |
| blaOXA-486, blaPDC | 113 | 0.37 |
| blaOXA-1028, blaPDC-3 | 111 | 0.36 |
| blaOXA-906, blaPDC-59 | 111 | 0.36 |
| blaOXA-494, blaPDC-71 | 108 | 0.35 |
| blaOXA-1033, blaPDC-3 | 105 | 0.34 |
| blaOXA-914, blaPDC-5 | 104 | 0.34 |
Carbapenemases
Unlike Acinetobacter baumannii, the intrinsic blaOXA alleles of P. aeruginosa are not typically associated with carbapenemase activity. Acquired carbapenemase genes, however, are present in 5,449 isolates (17.9%). The most common families of acquired carbapenemases are blaVIM (2,454 isolates, 8.1%), blaIMP (917 isolates, 3.0%), blaKPC (887 isolates, 2.9%), blaNDM (718 isolates, 2.4%), and blaGES (467 isolates, 1.5%), with seven additional gene families and two acquired blaOXA families found across 165 isolates (Table S15). A total of 11 distinct, assigned carbapenemase alleles occurred in 100 or more isolates each: 4 of 27 total encountered blaVIM carbapenemase alleles, 3 of 42 total encountered blaIMP carbapenemase alleles, 2 of 7 total encountered blaGES carbapenemase alleles, 1 of 6 total encountered blaKPC carbapenemase alleles, and 1 of 4 total encountered blaNDM carbapenemase alleles (Table S15).
Concerningly, metallo-β-lactamase genes with carbapenemase activity are present in 4,142 isolates (13.6%), many of which are likely to provide resistance to large swaths of the β-lactam armamentarium. The most common metallo-β-lactamases encountered were blaVIM-2 (1,703 isolates, 5.6%), blaNDM-1 (714 isolates, 2.3%), blaIMP-7 (227 isolates, 0.7%), blaVIM-1 (214 isolates, 0.7%), blaVIM-80 (209 isolates, 0.7%), blaIMP-1 (181 isolates, 0.6%), blaVIM-4 (163 isolates, 0.5%), and blaIMP-12 (103 isolates, 0.3%) (Table S15).
Alarmingly, carbapenemase genes appear to have become more common over time, increasing from 8.4% of isolates collected in 2000–2012 to 39.8% of isolates collected in 2022–2024 (Table S16). The overall proportion of isolates containing carbapenemase genes ranges from 11.5% in Oceania to 42.7% in South America (Table S17; Fig. 4). Interestingly, isolates without collection date or collection location metadata are far less likely to encode carbapenemase alleles (7.7% and 4.1%, respectively), but the reasons behind this are not clear.
Fig 4.
Proportion of isolates encoding carbapenemase genes by region. A total of 5,606 isolates (19.6%) encode carbapenemase genes.
Breadth of allelic diversity
As isolates were not collected under a uniform protocol, both over- and under-representation of alleles are likely. To address this, we “deduplicated” and examined the binary presence of distinct alleles within each group, which helps to mitigate the impact of successful clones and outbreaks and amplify the signal of widespread but low-frequency alleles.
MLST and pathogen detection SNP cluster frequency
Multilocus sequence typing (MLST) (17) and whole genome multilocus sequence typing (wgMLST) (18) are common approaches to determining the relatedness of bacterial isolates. The former utilizes a curated set of typically seven slowly evolving housekeeping genes to determine allele variants and assign a sequence type (ST), while the latter utilizes a much larger set of genes extracted from WGS data to determine the relatedness of isolates (19). NCBI Pathogen Detection generates wgMLST-based nearest-neighbor clusters of very closely related isolates, using a 25-allele cutoff (https://www.ncbi.nlm.nih.gov/pathogens/pathogens_help/#data-processing-clustering), which are given accessions starting with “PDS” (standing for Pathogen Detection Single nucleotide polymorphism [SNP]) and herein referred to as PDS clusters. Note that PDS clusters can change over time and are not archived, so the exact accessions mentioned below pertain only to the specific snapshot of data used in this analysis. The general patterns and trends should hold true across releases.
We found 1,863 distinct STs by MLST. The most common were ST235 (2,118 isolates, 7.0%) and ST111 (1,142 isolates, 3.8%), and nine others contain 2.0% or more of isolates: ST253, ST244, ST308, ST395, ST17, ST274, ST179, ST357, and ST155 (Table S18). By PDS, the data set contains 3,288 distinct clusters, of which PDS000112090.129 is the most common (640 isolates, 2.1%) with two other clusters occurring in 1.0% or more of isolates: PDS000095640.35 and PDS000012474.54. (Table S19). An additional 12,153 isolates were not assigned to a cluster, suggesting they are not closely related to any other isolate.
Changes in the frequency of STs were observed over time and geography. Temporally, ST235 is among the 2 most common across all periods, ST111 is 2nd through 6th most common across all periods, and ST253 is between 5th and 11th most common across all periods, none of which appear to be trending. Several STs exhibit dramatic variability between periods, perhaps associated with outbreaks or changes in the population structure over time: ST146 was the 9th most common in 2000–2012, 18th in 2013–2016, no higher than 28th in any other period, and was not observed in 2022–2024; ST316 was the 2nd most common in 2020–2021 but no higher than 27th most common in any other period, and ST1076 was the 8th most common in 2000–2012 but no higher than 21st most common in any other period, and ST1203 was the 3rd most common in 2022–2024 but no higher than 36th in any other period and not observed in 2000–2012 (Table S20). Regionally, ST235 is among the 2 most common in all regions except Sub-Saharan Africa (5th most common), ST111 is 2nd to 5th most common across six regions, but 16th in South Asia and 21st in East and Southeast Asia, and ST253 is between 4th and 15th most common across all regions. Among the most regionally variable STs, ST463 is the 2nd most common in East and Southeast Asia but no higher than 16th in any other region and not found in four regions; ST357 is the 26th most common in North America but 1st to 9th most common in all other regions; ST395 is the third most common in Europe and Central Asia but no higher than 12th most common in any other region and not found in two regions; and ST649 is the sixth most common in Oceania but no more than 16th most common in any other region and not found in four regions (Table S21).
An overview of allele frequency when the data are examined for the presence of individual alleles by ST and PDS clusters to better account for closely related isolates is provided in the supplemental text and Tables S22 and 23.
MLST and PDS association with intrinsic alleles
Given that sequence types and PDS clusters both serve to group related alleles, the question of whether certain STs and clusters are associated with specific alleles or whether certain alleles are associated with specific clusters (as might be expected if alleles are primarily disseminated in a clonal manner) naturally arises.
Looking at MLST, 1,574 (89.4%) of 1,761 STs associated with an assigned blaPDC allele correspond to a single allele, 138 (7.8%) to two alleles, 26 (1.5%) to three alleles, and 23 (1.3%) to four or more alleles, including ST235 (2,110 isolates corresponding to nine distinct alleles) and ST308 (754 isolates corresponding to 10 distinct alleles). Interestingly, intrinsic blaOXA alleles are more strongly associated with STs—1,605 (96.3%) of 1,667 STs associated with assigned intrinsic blaOXA alleles correspond to a single allele, 57 (including ST235) to two alleles, and 5 to three alleles (Table S24). Compared to STs, PDS clusters associated with assigned blaPDC alleles generally correspond to fewer alleles, with the vast majority of clusters (including the 10 most common, encompassing a combined 2,840 isolates) encoding just a single blaPDC allele, 46 clusters (1.4%) encoding two alleles, 1 cluster encoding three alleles, and 1 cluster (PDS000064916.3, consisting of 24 isolates) encoding four alleles. Clusters associated with assigned blaOXA alleles again correspond to far fewer alleles, with 3,141 (99.7%) of 3,150 clusters corresponding to a single allele and 9 clusters (0.3%) corresponding to two alleles (Table S25).
While STs and clusters are often associated with low numbers of distinct alleles, the opposite is not true. Among the 245 assigned blaPDC alleles associated with defined STs, the most widespread are blaPDC-3 (542 STs, 29.1%), blaPDC-5 (367 STs, 19.7%), blaPDC-9 (113 STs, 6.1%), and blaPDC-1 (105 STs, 5.6%), with a total of 52 alleles found in five or more STs each and only 120 alleles (49.0%) with a single ST. Of the 61 assigned blaOXA alleles associated with defined STs, the most widespread are blaOXA-494 (304 STs, 16.3%), blaOXA-50 (268 STs, 14.4%), and blaOXA-486 (232 STs, 12.5%), with a total of 40 alleles (65.6%) corresponding to five or more STs and only 10 alleles (16.4%) to single ST each (Table S26). Similar patterns emerge with clusters: of the 156 assigned blaPDC alleles associated with clusters, the most widespread are blaPDC-3 (622 clusters, 18.9%), blaPDC-5 (492 clusters, 15.0%), blaPDC-8 (350 clusters, 10.6%), blaPDC-1 (248 clusters, 7.5%), blaPDC-35 (244 clusters, 7.4%), blaPDC-19a (127 clusters, 3.9%), blaPDC-16 (125 clusters, 3.8%), and blaPDC-24 (102 clusters, 3.1%), with a total of 46 alleles (29.5%) corresponding to five or more clusters and 70 alleles (44.9%) to a single cluster. Among the 52 assigned, intrinsic blaOXA alleles associated with clusters, the most widespread are blaOXA-488 (528 clusters, 16.1%), blaOXA-50 (429 clusters, 13.0%), blaOXA-494 (400 clusters, 12.2%), blaOXA-486 (372 clusters, 11.3%), blaOXA-396 (234 clusters, 7.1%), blaOXA-395 (192 clusters, 5.8%), and blaOXA-847 (159 clusters, 4.8%), with a total of 31 alleles (59.6%) corresponding to five or more clusters and just 9 alleles (17.3%) to a single cluster (Table S27).
MLST and PDS granularity
Unsurprisingly, the higher resolution of wgMLST means PDS clusters provide more granularity than STs, resulting in a more diverse (and likely more representative) deduplicated data set. This difference in granularity is exemplified by the tendency of a single ST to correspond to many clusters, while the opposite is generally not true. Notably, ST235 (encompassing 7.0% of isolates) corresponds to 244 PDS clusters, and ST244 (encompassing 2.9% of isolates) corresponds to 122 PDS clusters, while 59 additional STs correspond to 10 or more clusters each, and 1,256 STs (67.4%) correspond to only a single cluster. Conversely, just 1 of 3,288 PDS clusters corresponds to three STs, 13 (0.5%) correspond to two STs, and 3,151 (95.8%) correspond to only a single ST (Table S28).
BioProject
BioProjects link isolates collected as part of the same research projects or studies, suggesting that it may be possible for isolates within them to be either closely related or selected based on similar phenotypic criteria, potentially impacting the results of the analysis. A total of 1,275 BioProjects were included in the data set, the largest of which vary greatly by region, ranging from 2,083 isolates (North America) to 69 isolates (Middle East and North Africa). Interestingly, the median number of isolates belonging to a BioProject is one or two across all regions, suggesting the results are not overrun with large BioProjects in general (Table S29).
Examining the presence of alleles by BioProject has minimal impact on overall allele frequency. The two most common blaPDC alleles, blaPDC-3 (455 BioProjects) and blaPDC-5 (377 BioProjects), remain unchanged, and none of the 10 most common alleles shift by more than one rank. Among intrinsic blaOXA, the two most common alleles, blaOXA-488 (470 BioProjects) and blaOXA-50 (455 BioProjects), remain unchanged, and none of the 10 most common alleles shift by more than three ranks (Table S30).
Comparing Pseudomonas alleles
Combining the sequences of assigned alleles present in the Reference Gene Catalog with those of unassigned alleles present in the data set provides a substantial snapshot into the diversity of β-lactamase alleles in P. aeruginosa.
Amino acid conservation
Conservation among blaPDC variants
Determining amino acid conservation among the distinct blaPDC alleles present in the data set and using the results to color code an experimentally determined crystal structure provide a residue-by-residue overview of conservation in PDC (Fig. 5A). For reference, Fig. 5B highlights important regions and residues of these class C β-lactamases, including the catalytic serine (S64), the general base lysine (K67), the YSN motif (containing the general base Y150), the Ω-loop (residues 188–221), the R2-loop (residues 289–307), and the KTG motif (containing K315) (20–22).
Fig 5.
Residue-by-residue conservation of blaPDC alleles. (A) Color-coded conservation of amino acid residues. (B) Important regions and motifs of class C β-lactamases, including the catalytic serine (S64; red), general base lysine (K67; purple), general base tyrosine (Y150, as part of the YSN motif; yellow), the Ω-loop (blue), the R2-loop (forest green), and the KTG motif (golden orange). Visualized on the crystal structure of PDC-1 using PDB ID: 4GZB (23).
As anticipated, the essential class C β-lactamase catalytic residues and characteristic motifs (S64-V65-S66-K67, Y150-S151-N152, and K315-T316-G317 by the [SANC] amino acid numbering convention) (20, 21) are highly conserved, as are several important residues in the C3/C4 carboxylate recognition region. The core portions of α-helices are also generally well conserved. Regions of lower conservation are scattered throughout, including portions of both the Ω-loop and R2-loop, regions known to play major roles in determining substrate specificity and the ability of variant enzymes to hydrolyze newer, larger, or otherwise non-native β-lactams and to overcome β-lactamase inhibitors (20, 22).
Conservation among blaOXA variants
Similarly analyzing blaOXA-50 family members provides a residue-by-residue overview of amino acid conservation in the blaOXA-50 family (Fig. 6A). For reference, Fig. 6B highlights important regions and residues of these class D β-lactamases, including the catalytic serine (S70), general base lysine (K73), the SXV motif, the (Y/F)GN motif, the K(T/S)G motif, and the Ω-loop (residues 160–169 by the OXA-66 crystal structure; approximately residues 152–167 by DBL [class D β-lactamase] numbering, some of which are not present in OXA-66) (21, 24, 25).
Fig 6.
Residue-by-residue conservation of blaOXA-50 family alleles. (A) Color-coded conservation of amino acid residues. (B) Important regions and motifs of class D β-lactamases, including the catalytic serine (S70; red), general base lysine (K73; purple), the SXV motif (lime green), the (Y/F)GN motif (yellow), the K(T/S)G motif (golden orange), and the Ω-loop (blue). Visualized on a SWISS-MODEL homology model of OXA-488 based on the crystal structure of OXA-66.
As anticipated, the essential class D β-lactamase catalytic residues (or biochemical properties of the residues) and characteristic motifs (S70-T71-F72/Y72-K73, S118-X119-V120/I120, Y144-G145-N146, and K216-T217/S217-G218) (21, 26) are highly conserved. The apparently lower conservation of I120 is the result of several amino acids with hydrophobic sidechains being present in that position, likely maintaining many of the properties even as the position becomes somewhat variable.
Protein sequence identity among blaPDC variants
When aligning and examining predicted PDC protein sequences, an interesting observation emerges: the mature proteins of 520 assigned alleles have 98% or higher identity to PDC-3 (the most common PDC variant), 43 have 96%–98% identity, and 5 have 93%–96% identity, while an outlier group of 13 alleles have only 88%–89% identity to PDC-3 but share at least 98% identity with each other, differing by no more than six residues. While the significance of this apparent bifurcation of PDC variants is not immediately clear, one variant (PDC-33) is associated with an isolate previously described as a “taxonomic outlier Pseudomonas aeruginosa” (27). Further genetic and biochemical investigation of these variants (and the isolates producing them) is warranted to determine the origin and implications of these outliers.
DISCUSSION
Insights into the resistome
Many β-lactamase alleles remain unstudied. While P. aeruginosa is a generally well-studied organism, the ratio of distinct assigned to unassigned alleles is surprising, with a total of 622 distinct, unassigned blaPDC alleles present in the data set compared to only 254 distinct, assigned alleles, revealing substantial unstudied diversity (as allele assignments are requested by researchers submitting them [28], unassigned alleles are generally unstudied alleles). Interestingly, we only encountered 43.6% of the 582 total assigned blaPDC alleles present in the Reference Gene Catalog, further underscoring the rarity of many distinct alleles. This represents a large reservoir of potentially concerning alleles, and understanding and characterizing those with the most resistant phenotypes is an essential area of β-lactamase research.
Carbapenemase prevalence is high and increasing. Carbapenems are broad-spectrum β-lactams that have historically been used as an antibiotic of “last resort” for the treatment of complicated and multidrug-resistant infections (29). As a major concern in MDR P. aeruginosa infections, understanding the diversity and prevalence of these enzymes is crucial to overcoming these infections. Disturbingly, the presence of carbapenemases appears to have greatly increased (or carbapenemase-containing isolates have become a bigger proportion of isolates being sequenced and deposited into NCBI databases) over the past two decades. Compared to isolates collected from 2000 to 2012, carbapenemase prevalence is nearly fivefold higher for P. aeruginosa isolates collected from 2022 to 2024. Similar observations have been made about increasing carbapenem resistance (30), but further work is needed to better understand the cause of this increase and determine how to best address it.
The distribution of β-lactamase alleles is heavily skewed, and most are uncommon. In 30,452 P. aeruginosa isolates, we only encountered 254 of 582 assigned blaPDC alleles (43.6%). Of these, the majority are extremely uncommon, with 171 blaPDC alleles (67.3%) found in 10 or fewer isolates each. Conversely, just five blaPDC alleles (blaPDC-3, blaPDC-5, blaPDC-8, blaPDC-1, and blaPDC-35) accounted for nearly 60% of all blaPDC alleles present.
While investigating the causes of this imbalance remains the territory of future studies, two potential hypotheses are (i) the most common alleles being more ancestral versions of their respective genes, and (ii) purifying selection favoring these variants over others of similar phenotypes. Interestingly, their presence in highly successful clones is likely not sufficient to explain this phenomenon as analysis by both MLST and PDS clusters suggests that they are widespread among very dissimilar isolates.
Breadth versus depth of allelic diversity
Due to factors including the characteristics of the studies providing sequencing data, outbreaks, and “successful clones,” among others, some closely related groups of isolates are likely overrepresented in the data set. While this may reflect larger trends in the overall population, it runs the risk of masking emerging alleles and deemphasizes alleles arising and disseminating in multiple lineages. In a sense, the “depth” (allele frequency) of diversity overwhelms the “breadth” (number of distinct alleles) of diversity, increasing the odds of important alleles being missed. By deduplicating allele frequencies either by closely related isolates (e.g., MLST and PDS clusters) or isolates collected for a common purpose, often with shared and carefully designed inclusion criteria (BioProject), we can form a better picture of the breadth of allelic diversity without the overshadowing of depth. Interestingly, deduplication by ST, PDS clusters and unclustered isolates, and BioProject all led to generally very minor changes for most alleles, suggesting the overall data set is not greatly skewed as a result and that our analysis is a reasonable approach.
Overall, sequence types are more closely associated with intrinsic blaOXA alleles than blaPDC alleles, with the latter demonstrating more intra-ST variability (the most widespread ST in the data set, ST235, is associated with nine blaPDC alleles but only two intrinsic blaOXA alleles across 2,118 isolates). While both sequence types and clusters are typically associated with a small number of alleles, the opposite is not true—many alleles can be found in 10s to 100s of different STs and clusters (the most widespread blaPDC allele, blaPDC-3, is associated with 542 distinct STs and 622 distinct clusters).
Additionally, PDS clusters provide a far more granular picture of isolates than MLST sequence types. The most common sequence type in P. aeruginosa corresponds to 244 PDS clusters, while the most common cluster corresponds to just three sequence types. This is not entirely surprising given the much higher resolution associated with the wgMLST origin of PDS clusters, the full extent of this difference is interesting and demonstrates a potential downside of traditional MLST.
Amino acid conservation
Comparing the sequences of all assigned blaPDC and blaOXA-50 family alleles and corresponding unassigned alleles appearing in this data set, several residues appear less conserved than expected, most notably the catalytic S64 of PDC. Upon closer examination, however, it becomes apparent that the residue is nearly universally conserved with just two variants—PDC-47 and an unassigned PDC— harboring S64L and S64P substitutions, respectively. PDC-47 has only been reported in a single isolate, which the authors expected to be nonfunctional (31). Likewise, we hypothesize that the lack of an essential catalytic residue renders both variants incapable of hydrolysis and unable to contribute to resistance in strains harboring them.
Comparisons to surveillance studies
The Antibiotic Resistance Leadership Group’s Prospective Observational Pseudomonas (POP) study examined CRPa isolates collected from 972 patients across 10 countries in 2018 and 2019 (8). As the POP sequences are available in NCBI databases, 712 isolates belonging to the POP BioProject (PRJNA824880) are included in our study. Collectively, they represent 65 (2.7%) of 2,385 isolates from 2018 and 648 (29.4%) of 2,205 isolates from 2019. These isolates have been excluded from our results for the following comparison and discussion.
POP reported carbapenemase genes in 22% of isolates (8), slightly higher than our findings of 18.5% in 2018 and 2019, but this difference is in line with expectations considering their focus on CRPa. The five most common carbapenemase gene families (blaVIM, blaNDM, blaKPC, blaIMP, and blaGES) were the same in POP and our present study, but they found blaKPC to be the most common (8). Although defining regions differently, the region containing South America has the highest prevalence of carbapenemase genes in both studies, and the region containing North America is among the lowest (8). We found the most common sequence types to be ST111, ST308, ST463, and ST235, while the POP study found them to be ST235, ST111, ST463, and ST308 ranked differently (8). Overall, many general trends hold true between our work and the POP study, suggesting our methodology is successfully capturing allelic diversity in P. aeruginosa.
Limitations
Although the data set was processed to remove incomplete sequences and “lower-quality” allele calls, some (particularly unassigned) alleles may be inaccurate as the result of sequencing errors. As a genetic and bioinformatic analysis, we did not evaluate protein expression (including derepression and overexpression that can lead to high-level β-lactam resistance in P. aeruginosa [32]) or stability, nor the resistance phenotypes of the isolates encoding these alleles. Further studies of individual isolates and enzymes would help to address phenotypic and protein-related questions.
Additionally, while providing a snapshot of allelic diversity across time and place, we have not conducted a surveillance study. In comparison to the more rigid inclusion criteria of a surveillance study, some selection bias exists. The isolates sequenced, submitted, and ultimately used in this analysis were selected based on the needs and interests of the original researchers and studies providing sequencing data, as opposed to a protocol collecting consecutive isolates directly from clinical laboratories in an unbiased manner. Unfortunately, this effect is an intrinsic property of the data set and cannot be readily quantified but is reduced by including as large a sample of isolates as possible.
Conclusions
In exploring the frequency and distribution of β-lactamase alleles, we have determined important alleles and combinations that warrant coverage by future β-lactams and inhibitors targeting P. aeruginosa and that clinicians should be aware of. While this data set is not ideal for examining evolution or selective pressure, we speculate that the existence of a small number of extremely successful enzymes suggests an evolutionary advantage worthy of further study.
With rapidly growing collections of whole genome sequenced bacteria, the NCBI Pathogen Detection Project provides a rich and diverse pool of data on resistance alleles. Querying these data enables exploration of the distribution and diversity of β-lactamase alleles present in P. aeruginosa (and, more broadly, the antimicrobial resistance, virulence, and stress response genes in any of the 98 organism groups included in the database). As this data set is diverse in terms of time, geography, and species, we suggest this approach provides useful snapshots of allelic diversity and may serve as an alternative to more traditional, expensive, and time-consuming surveillance studies for examining the distribution and diversity of AMR and related alleles when the epidemiological precision of such a study is not required.
Finally, there remains a large pool of β-lactamase allelic diversity in P. aeruginosa, which has largely gone unstudied and unrecognized. Given the high level of plasticity and ease of evolution of β-lactamase enzymes, one can speculate that variants that have not yet faced the right selective pressures could easily spread causing future outbreaks. Ultimately, further microbiological, biochemical, and evolutionary biology studies are needed to truly understand the meaning of this diversity. We assert that understanding the scope and distribution of this diversity is crucial to prioritizing future research and development efforts to focus on the most common and widespread resistance alleles of today and perhaps begin to predict what the future may hold.
MATERIALS AND METHODS
Data sources and processing
Data were obtained from databases maintained by the National Center for Biotechnology Information (NCBI), primarily curated as part of the Pathogen Detection Program (https://www.ncbi.nlm.nih.gov/pathogens).
Microbial Browser for Identification of Genetic and Genomic Elements (MicroBIGG-E) (33) data were queried on 9 September 2024 using the Google Cloud Platform interface (https://www.ncbi.nlm.nih.gov/pathogens/docs/microbigge_gcp/) with the query “SELECT * FROM `ncbi-pathogen-detect.pdbrowser.microbigge` WHERE (`scientific_name` like 'Pseudomonas aeruginosa') AND (`element_symbol` like '%bla%').”
Reference Gene Catalog release “2024-07-22.1” (34) was downloaded from the NCBI FTP server (https://ftp.ncbi.nlm.nih.gov/pathogen/Antimicrobial_resistance/Data/2024-07-22.1/ReferenceGeneCatalog.txt) and used to supplement allele and OXA family names as needed.
Identical Protein Group (IPG) accessions were matched to all unassigned alleles and each distinct assigned allele. These accessions were used to group distinct unassigned alleles and update missing allele assignments for alleles added to the Reference Gene Catalog after isolates were processed. IPG identifiers were determined by querying protein accession numbers from MicroBIGG-E using the NCBI Entrez Eutils interface on 9 September 2024 via the REntrez R package.
For purposes of this analysis, alleles are designated as either “assigned” or “unassigned” based on the presence of a formally assigned allele designation. “Assigned” alleles refer to those with a formal designation assigned by NCBI (or Institut Pasteur for blaLEN, blaOKP-A, blaOKP-B, and blaOXY) and an entry in the Reference Gene Catalog, while “unassigned” alleles refer to those lacking a formal designation and not appearing in the Reference Gene Catalog. Assigned alleles are represented in Pathogen Detection databases with a gene name and number (e.g., blaPDC-3), while unassigned alleles are represented only by a gene name (e.g., blaPDC).
To help account for potential sequencing errors and differences in sequence quality (resulting from the use of a large, publicly available data set rather than sequencing and verifying isolates directly), we removed lower-quality alleles (partial alleles with less than 100% coverage, unassigned alleles with less than 90% identity to the closest reference allele, partial alleles crossing contig boundaries, and mistranslations/internal stop codons) from the data set as a final processing step. We also removed alleles identified only by hidden Markov models in AMRFinderPlus (35, 36), as they typically do not contain sufficient detail for a complete analysis. Finally, we removed any isolates not encoding both a blaPDC allele and a blaOXA-50 family member allele, as these are important, intrinsic genes expected in P. aeruginosa.
The compiled, processed, and filtered data set used in our analysis has been uploaded to Zenodo as Table S1, available from https://doi.org/10.5281/zenodo.13917360.
Data analysis
The analysis was performed in RStudio version 2024.04.2 Build 764 using R version 4.4.1.
MLST determinations were made using FastMLST version 0.0.16 (37) with sequences downloaded from NCBI (38). The latest allele and sequence type definitions were downloaded from PubMLST (39) on 10 September 2024.
Original code for downloading, cleaning, analysis, and graphic generation has been prepared as the R package “pdallele.” The exact version used herein is available from https://doi.org/10.5281/zenodo.13917356, and the most current version is available from https://github.com/armack/pdallele/.
Data notes
The source data use the International Nucleotide Sequence Database Collaboration “country” controlled vocabulary (40), which includes oceans, seas, and several territories and islands as “countries.” As this differs from the definition of “country” used in this analysis, we have processed these country names using the “countrycode” R package, resulting in the dropping of geographic information for ocean/sea locations and some shared or disputed islands from the analysis. Regions (Table S31) are defined based on groupings of United Nations geoscheme subregions (41). We chose to utilize these non-standard regional groupings primarily to provide for the separation of North and South America (which most existing United Nations and World Health Organization regions do not).
Pathogen Detection databases do not explicitly report chromosomal and non-chromosomal alleles for most isolates (i.e., those not assembled at the “complete-genome” level), and as most assemblies do not have chromosomal versus plasmid contigs annotated as such, we assume that blaPDC and blaOXA-50 family alleles are chromosomal, while all other β-lactamase alleles are acquired.
All gene and allele names, blaOXA families, and carbapenemase designations are used and discussed as provided in the Pathogen Detection databases.
Alignments and amino acid conservation
Sequences of all assigned blaPDC and blaOXA-50 family alleles appearing in the Reference Gene Catalog and corresponding unassigned alleles appearing in the MicroBIGG-E data set were downloaded in FASTA format using Identical Protein Group identifiers. Alignments were determined using the MUSCLE3 algorithm (42) with default settings as bundled in UniPro UGENE version 50.0 (43). Amino acid conservation was determined using AL2CO (44) with the default independent counts frequency estimation and entropy-based conservation measure settings as provided in UCSF ChimeraX version 1.8 (45) and used to color code the 4GZB crystal structure of PDC-3 (23) from the RCSB Protein Data Bank (46). Due to the lack of an available crystal structure of an OXA-50 family member, a SWISS-MODEL (47) homology model of OXA-488 was generated based on the 6T1H crystal structure of OXA-66 and used to represent the family.
ACKNOWLEDGMENTS
The research reported herein was supported in part by funds from the National Institute of Allergy and Infectious Diseases of the National Institutes of Health under Award Number R01AI063517 to R.A.B. The work of D.H.H., M.F., W.K., and A.B.P. was supported by the National Center for Biotechnology Information of the National Library of Medicine (NLM) and the National Institute of Allergy and Infectious Diseases, National Institutes of Health.
We thank Rachel A. Powers, PhD, for critically reading the manuscript.
The content is solely the responsibility of the authors and does not necessarily represent the official views of the Department of Veterans Affairs or the National Institutes of Health.
A.R.M., R.A.B., A.M.H., M.F.M., and M.A.T. conceptualized the study; M.F., D.H.H., W.K., and A.B.P. curated the data; A.R.M. carried out formal analysis and investigation, designed the methodology, provided software, visualized the study, and wrote the original draft; R.A.B. acquired funding; R.A.B., A.M.H., and W.K. administrated the project; R.A.B. and A.M.H. supervised the study; A.R.M., A.M.H., M.F.M., M.A.T., M.F., D.H.H., W.K., A.B.P., and R.A.B. reviewed and edited the manuscript.
Contributor Information
Robert A. Bonomo, Email: Robert.Bonomo@va.gov.
Laurent Poirel, University of Fribourg, Fribourg, Switzerland.
SUPPLEMENTAL MATERIAL
The following material is available online at https://doi.org/10.1128/aac.00785-24.
Regions used in the analysis and isolate counts for each region.
Additional analysis using allele presence data.
Tables S2 to S31.
ASM does not own the copyrights to Supplemental Material that may be linked to, or accessed through, an article. The authors have granted ASM a non-exclusive, world-wide license to publish the Supplemental Material files. Please contact the corresponding author directly for reuse.
REFERENCES
- 1. Reynolds D, Kollef M. 2021. The epidemiology and pathogenesis and treatment of Pseudomonas aeruginosa infections: an update. Drugs (Abingdon Engl) 81:2117–2131. doi: 10.1007/s40265-021-01635-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Pang Z, Raudonis R, Glick BR, Lin T-J, Cheng Z. 2019. Antibiotic resistance in Pseudomonas aeruginosa: mechanisms and alternative therapeutic strategies. Biotechnol Adv 37:177–192. doi: 10.1016/j.biotechadv.2018.11.013 [DOI] [PubMed] [Google Scholar]
- 3. World Health Organization . 2024. WHO bacterial priority pathogens list, 2024: bacterial pathogens of public health importance to guide research, development and strategies to prevent and control antimicrobial resistance. World Health Organization, Geneva, Switzerland. [Google Scholar]
- 4. Centers for Disease Control and Prevention . 2019. Antibiotic resistance threats in the United States, 2019. Centers for Disease Control and Prevention. https://www.cdc.gov/antimicrobial-resistance/data-research/threats/index.html. [Google Scholar]
- 5. Centers for Disease Control and Prevention . 2022. COVID-19: U.S. impact on antimicrobial resistance, special report 2022. National Center for Emerging and Zoonotic Infectious Diseases. https://stacks.cdc.gov/view/cdc/119025. [Google Scholar]
- 6. Shortridge D, Gales AC, Streit JM, Huband MD, Tsakris A, Jones RN. 2019. Geographic and temporal patterns of antimicrobial resistance in Pseudomonas aeruginosa over 20 years from the SENTRY antimicrobial surveillance program, 1997-2016. Open Forum Infect Dis 6:S63–S68. doi: 10.1093/ofid/ofy343 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Lee Y-L, Ko W-C, Hsueh P-R. 2022. Geographic patterns of carbapenem-resistant Pseudomonas aeruginosa in the Asia-Pacific region: results from the antimicrobial testing leadership and surveillance (ATLAS) Program, 2015-2019. Antimicrob Agents Chemother 66:e0200021. doi: 10.1128/AAC.02000-21 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Reyes J, Komarow L, Chen L, Ge L, Hanson BM, Cober E, Herc E, Alenazi T, Kaye KS, Garcia-Diaz J, et al. 2023. Global epidemiology and clinical outcomes of carbapenem-resistant Pseudomonas aeruginosa and associated carbapenemases (POP): a prospective cohort study. Lancet Microbe 4:e159–e170. doi: 10.1016/S2666-5247(22)00329-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Karlowsky JA, Lob SH, DeRyke CA, Hilbert DW, Wong MT, Young K, Siddiqui F, Motyl MR, Sahm DF. 2022. In vitro activity of ceftolozane-tazobactam, imipenem-relebactam, ceftazidime-avibactam, and comparators against Pseudomonas aeruginosa isolates collected in United States hospitals according to results from the SMART surveillance program, 2018 to 2020. Antimicrob Agents Chemother 66:e0018922. doi: 10.1128/aac.00189-22 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Nichols WW, de Jonge BLM, Kazmierczak KM, Karlowsky JA, Sahm DF. 2016. In Vitro Susceptibility of Global Surveillance Isolates of Pseudomonas aeruginosa to Ceftazidime-Avibactam (INFORM 2012 to 2014). Antimicrob Agents Chemother 60:4743–4749. doi: 10.1128/AAC.00220-16 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Dunphy LJ, Kolling GL, Jenior ML, Carroll J, Attai AE, Farnoud F, Mathers AJ, Hughes MA, Papin JA. 2021. Multidimensional clinical surveillance of Pseudomonas aeruginosa reveals complex relationships between isolate source, morphology, and antimicrobial resistance. mSphere 6:e0039321. doi: 10.1128/mSphere.00393-21 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Sloot R, Nsonwu O, Chudasama D, Rooney G, Pearson C, Choi H, Mason E, Springer A, Gerver S, Brown C, Hope R. 2022. Rising rates of hospital-onset Klebsiella spp. and Pseudomonas aeruginosa bacteraemia in NHS acute trusts in England: a review of national surveillance data, August 2020-February 2021. J Hosp Infect 119:175–181. doi: 10.1016/j.jhin.2021.08.027 [DOI] [PubMed] [Google Scholar]
- 13. Antibiotic Resistance and Pathogenicity of Bacterial Infections Group – IdISBa . 2018. Pseudomonas aeruginosa derived cephalosporinase (PDC) database. https://arpbigidisba.com/pseudomonas-aeruginosa-derived-cephalosporinase-pdc-database/.
- 14. Lebreton F, Snesrud E, Hall L, Mills E, Galac M, Stam J, Ong A, Maybank R, Kwak YI, Johnson S, Julius M, Ly M, Swierczewski B, Waterman PE, Hinkle M, Jones A, Lesho E, Bennett JW, McGann P. 2021. A panel of diverse Pseudomonas aeruginosa clinical isolates for research and development. JAC Antimicrob Resist 3:dlab179. doi: 10.1093/jacamr/dlab179 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. National Library of Medicine (US), National Center for Biotechnology Information . 2021. Pathogen detection help document. Available from: https://www.ncbi.nlm.nih.gov/pathogens/pathogens_help/. Retrieved 13 Mar 2023.
- 16. O’Leary NA, Wright MW, Brister JR, Ciufo S, Haddad D, McVeigh R, Rajput B, Robbertse B, Smith-White B, Ako-Adjei D, et al. 2016. Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Res 44:D733–45. doi: 10.1093/nar/gkv1189 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Maiden MC, Bygraves JA, Feil E, Morelli G, Russell JE, Urwin R, Zhang Q, Zhou J, Zurth K, Caugant DA, Feavers IM, Achtman M, Spratt BG. 1998. Multilocus sequence typing: a portable approach to the identification of clones within populations of pathogenic microorganisms. Proc Natl Acad Sci U S A 95:3140–3145. doi: 10.1073/pnas.95.6.3140 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Sheppard SK, Jolley KA, Maiden MCJ. 2012. A gene-by-gene approach to bacterial population genomics: whole genome MLST of Campylobacter. Genes (Basel) 3:261–277. doi: 10.3390/genes3020261 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Maiden MCJ, Jansen van Rensburg MJ, Bray JE, Earle SG, Ford SA, Jolley KA, McCarthy ND. 2013. MLST revisited: the gene-by-gene approach to bacterial genomics. Nat Rev Microbiol 11:728–736. doi: 10.1038/nrmicro3093 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Mack AR, Barnes MD, Taracila MA, Hujer AM, Hujer KM, Cabot G, Feldgarden M, Haft DH, Klimke W, van den Akker F, et al. 2020. A standard numbering scheme for class C β-lactamases. Antimicrob Agents Chemother 64:e01841-19. doi: 10.1128/AAC.01841-19 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Bush K. 2013. The ABCD’s of β-lactamase nomenclature. J Infect Chemother 19:549–559. doi: 10.1007/s10156-013-0640-7 [DOI] [PubMed] [Google Scholar]
- 22. Jacoby GA. 2009. AmpC β-Lactamases. Clin Microbiol Rev 22:161–182. doi: 10.1128/CMR.00036-08 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Lahiri SD, Mangani S, Durand-Reville T, Benvenuti M, De Luca F, Sanyal G, Docquier J-D. 2013. Structural insight into potent broad-spectrum inhibition with reversible recyclization mechanism: avibactam in complex with CTX-M-15 and Pseudomonas aeruginosa AmpC β-lactamases. Antimicrob Agents Chemother 57:2496–2505. doi: 10.1128/AAC.02247-12 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Joris B, Ledent P, Dideberg O, Fonzé E, Lamotte-Brasseur J, Kelly JA, Ghuysen JM, Frère JM. 1991. Comparison of the sequences of class A beta-lactamases and of the secondary structure elements of penicillin-recognizing proteins. Antimicrob Agents Chemother 35:2294–2301. doi: 10.1128/AAC.35.11.2294 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Paetzel M, Danel F, de Castro L, Mosimann SC, Page MG, Strynadka NC. 2000. Crystal structure of the class D beta-lactamase OXA-10. Nat Struct Biol 7:918–925. doi: 10.1038/79688 [DOI] [PubMed] [Google Scholar]
- 26. Poirel L, Naas T, Nordmann P. 2010. Diversity, epidemiology, and genetics of class D β-lactamases. Antimicrob Agents Chemother 54:24–38. doi: 10.1128/AAC.01512-08 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Roy PH, Tetu SG, Larouche A, Elbourne L, Tremblay S, Ren Q, Dodson R, Harkins D, Shay R, Watkins K, Mahamoud Y, Paulsen IT. 2010. Complete genome sequence of the multiresistant taxonomic outlier Pseudomonas aeruginosa PA7. PLoS ONE 5:e8842. doi: 10.1371/journal.pone.0008842 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Bradford PA, Bonomo RA, Bush K, Carattoli A, Feldgarden M, Haft DH, Ishii Y, Jacoby GA, Klimke W, Palzkill T, Poirel L, Rossolini GM, Tamma PD, Arias CA. 2022. Consensus on β-lactamase nomenclature. Antimicrob Agents Chemother 66:e0033322. doi: 10.1128/aac.00333-22 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Papp-Wallace KM, Endimiani A, Taracila MA, Bonomo RA. 2011. Carbapenems: past, present, and future. Antimicrob Agents Chemother 55:4943–4960. doi: 10.1128/AAC.00296-11 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Wang M-G, Liu Z-Y, Liao X-P, Sun R-Y, Li R-B, Liu Y, Fang L-X, Sun J, Liu Y-H, Zhang R-M. 2021. Retrospective data insight into the global distribution of carbapenemase-producing Pseudomonas aeruginosa. Antibiotics (Basel) 10:548. doi: 10.3390/antibiotics10050548 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Lahiri SD, Johnstone MR, Ross PL, McLaughlin RE, Olivier NB, Alm RA. 2014. Avibactam and class C β-lactamases: mechanism of inhibition, conservation of the binding pocket, and implications for resistance. Antimicrob Agents Chemother 58:5704–5713. doi: 10.1128/AAC.03057-14 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Juan C, Moyá B, Pérez JL, Oliver A. 2006. Stepwise upregulation of the Pseudomonas aeruginosa chromosomal cephalosporinase conferring high-level beta-lactam resistance involves three AmpD homologues. Antimicrob Agents Chemother 50:1780–1787. doi: 10.1128/AAC.50.5.1780-1787.2006 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Feldgarden M, Brover V, Fedorov B, Haft DH, Prasad AB, Klimke W. 2022. Curation of the AMRFinderPlus databases: applications, functionality and impact. Microb Genom 8:mgen000832. doi: 10.1099/mgen.0.000832 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA. Reference Gene Catalog (Release 2024-07-22.1) (2024-07-22.1). 2024
- 35. Feldgarden M, Brover V, Haft DH, Prasad AB, Slotta DJ, Tolstoy I, Tyson GH, Zhao S, Hsu C-H, McDermott PF, Tadesse DA, Morales C, Simmons M, Tillman G, Wasilenko J, Folster JP, Klimke W. 2019. Validating the AMRFinder tool and resistance gene database by using antimicrobial resistance genotype-phenotype correlations in a collection of isolates. Antimicrob Agents Chemother 63:e00483-19. doi: 10.1128/AAC.00483-19 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Feldgarden M, Brover V, Gonzalez-Escalona N, Frye JG, Haendiges J, Haft DH, Hoffmann M, Pettengill JB, Prasad AB, Tillman GE, Tyson GH, Klimke W. 2021. AMRFinderPlus and the reference gene catalog facilitate examination of the genomic links among antimicrobial resistance, stress response, and virulence. Sci Rep 11:12728. doi: 10.1038/s41598-021-91456-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Guerrero-Araya E, Muñoz M, Rodríguez C, Paredes-Sabja D. 2021. FastMLST: a multi-core tool for multilocus sequence typing of draft genome assemblies. Bioinform Biol Insights 15:11779322211059238. doi: 10.1177/11779322211059238 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Sayers EW, Bolton EE, Brister JR, Canese K, Chan J, Comeau DC, Farrell CM, Feldgarden M, Fine AM, Funk K, et al. 2023. Database resources of the national center for biotechnology information in 2023. Nucleic Acids Res 51:D29–D38. doi: 10.1093/nar/gkac1032 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Jolley KA, Bray JE, Maiden MCJ. 2018. Open-access bacterial population genomics: BIGSdb software, the PubMLST.org website and their applications. Wellcome Open Res 3:124. doi: 10.12688/wellcomeopenres.14826.1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. International Nucleotide Sequence Database . 2014. Controlled vocabulary for /country qualifier. Available from: https://www.insdc.org/submitting-standards/country-qualifier-vocabulary. Retrieved 17 Oct 2023.
- 41. United Nations Statistics Division . 2021. Standard country or area codes for statistical use M49. United Nations, New York, NY, USA. [Google Scholar]
- 42. Edgar RC. 2004. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res 32:1792–1797. doi: 10.1093/nar/gkh340 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Okonechnikov K, Golosova O, Fursov M, the UGENE team . 2012. Unipro UGENE: a unified bioinformatics toolkit. Bioinformatics 28:1166–1167. doi: 10.1093/bioinformatics/bts091 [DOI] [PubMed] [Google Scholar]
- 44. Pei J, Grishin NV. 2001. AL2CO: calculation of positional conservation in a protein sequence alignment. Bioinformatics 17:700–712. doi: 10.1093/bioinformatics/17.8.700 [DOI] [PubMed] [Google Scholar]
- 45. Pettersen EF, Goddard TD, Huang CC, Meng EC, Couch GS, Croll TI, Morris JH, Ferrin TE. 2021. UCSF ChimeraX: structure visualization for researchers, educators, and developers. Protein Sci 30:70–82. doi: 10.1002/pro.3943 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46. Burley SK, Bhikadiya C, Bi C, Bittrich S, Chao H, Chen L, Craig PA, Crichlow GV, Dalenberg K, Duarte JM, et al. 2023. RCSB protein data bank (RCSB.org): delivery of experimentally-determined PDB structures alongside one million computed structure models of proteins from artificial intelligence/machine learning. Nucleic Acids Res 51:D488–D508. doi: 10.1093/nar/gkac1077 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Waterhouse A, Bertoni M, Bienert S, Studer G, Tauriello G, Gumienny R, Heer FT, de Beer TAP, Rempfer C, Bordoli L, Lepore R, Schwede T. 2018. SWISS-MODEL: homology modelling of protein structures and complexes. Nucleic Acids Res 46:W296–W303. doi: 10.1093/nar/gky427 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Regions used in the analysis and isolate counts for each region.
Additional analysis using allele presence data.
Tables S2 to S31.






