203 research outputs found

    Optimal contact map alignment of protein–protein interfaces

    Get PDF
    The long-standing problem of constructing protein structure alignments is of central importance in computational biology. The main goal is to provide an alignment of residue correspondences, in order to identify homologous residues across chains. A critical next step of this is the alignment of protein complexes and their interfaces. Here, we introduce the program CMAPi, a two-dimensional dynamic programming algorithm that, given a pair of protein complexes, optimally aligns the contact maps of their interfaces: it produces polynomial-time near-optimal alignments in the case of multiple complexes. We demonstrate the efficacy of our algorithm on complexes from PPI families listed in the SCOPPI database and from highly divergent cytokine families. In comparison to existing techniques, CMAPi generates more accurate alignments of interacting residues within families of interacting proteins, especially for sequences with low similarity. While previous methods that use an all-atom based representation of the interface have been successful, CMAPi's use of a contact map representation allows it to be more tolerant to conformational changes and thus to align more of the interaction surface. These improved interface alignments should enhance homology modeling and threading methods for predicting PPIs by providing a basis for generating template profiles for sequence–structure alignment

    Deriving amino acid contact potentials from their frequencies of occurence in proteins: a lattice model study

    Full text link
    The possibility of deriving the contact potentials between amino acids from their frequencies of occurence in proteins is discussed in evolutionary terms. This approach allows the use of traditional thermodynamics to describe such frequencies and, consequently, to develop a strategy to include in the calculations correlations due to the spatial proximity of the amino acids and to their overall tendency of being conserved in proteins. Making use of a lattice model to describe protein chains and defining a "true" potential, we test these strategies by selecting a database of folding model sequences, deriving the contact potentials from such sequences and comparing them with the "true" potential. Taking into account correlations allows for a markedly better prediction of the interaction potentials

    Numerous proteins with unique characteristics are degraded by the 26S proteasome following monoubiquitination

    Get PDF
    The "canonical" proteasomal degradation signal is a substrate-anchored polyubiquitin chain. However, a handful of proteins were shown to be targeted following monoubiquitination. In this study, we established-in both human and yeast cells-a systematic approach for the identification of monoubiquitination-dependent proteasomal substrates. The cellular wild-type polymerizable ubiquitin was replaced with ubiquitin that cannot form chains. Using proteomic analysis, we screened for substrates that are nevertheless degraded under these conditions compared with those that are stabilized, and therefore require polyubiquitination for their degradation. For randomly sampled representative substrates, we confirmed that their cellular stability is in agreement with our screening prediction. Importantly, the two groups display unique features: monoubiquitinated substrates are smaller than the polyubiquitinated ones, are enriched in specific pathways, and, in humans, are structurally less disordered. We suggest that monoubiquitination-dependent degradation is more widespread than assumed previously, and plays key roles in various cellular processes

    Structure of a putative NTP pyrophosphohydrolase: YP_001813558.1 from Exiguobacterium sibiricum 255-15.

    Get PDF
    The crystal structure of a putative NTPase, YP_001813558.1 from Exiguobacterium sibiricum 255-15 (PF09934, DUF2166) was determined to 1.78 Å resolution. YP_001813558.1 and its homologs (dimeric dUTPases, MazG proteins and HisE-encoded phosphoribosyl ATP pyrophosphohydrolases) form a superfamily of all-α-helical NTP pyrophosphatases. In dimeric dUTPase-like proteins, a central four-helix bundle forms the active site. However, in YP_001813558.1, an unexpected intertwined swapping of two of the helices that compose the conserved helix bundle results in a `linked dimer' that has not previously been observed for this family. Interestingly, despite this novel mode of dimerization, the metal-binding site for divalent cations, such as magnesium, that are essential for NTPase activity is still conserved. Furthermore, the active-site residues that are involved in sugar binding of the NTPs are also conserved when compared with other α-helical NTPases, but those that recognize the nucleotide bases are not conserved, suggesting a different substrate specificity

    Structure of the γ-D-glutamyl-L-diamino acid endopeptidase YkfC from Bacillus cereus in complex with L-Ala-γ-D-Glu: insights into substrate recognition by NlpC/P60 cysteine peptidases.

    Get PDF
    Dipeptidyl-peptidase VI from Bacillus sphaericus and YkfC from Bacillus subtilis have both previously been characterized as highly specific γ-D-glutamyl-L-diamino acid endopeptidases. The crystal structure of a YkfC ortholog from Bacillus cereus (BcYkfC) at 1.8 Å resolution revealed that it contains two N-terminal bacterial SH3 (SH3b) domains in addition to the C-terminal catalytic NlpC/P60 domain that is ubiquitous in the very large family of cell-wall-related cysteine peptidases. A bound reaction product (L-Ala-γ-D-Glu) enabled the identification of conserved sequence and structural signatures for recognition of L-Ala and γ-D-Glu and, therefore, provides a clear framework for understanding the substrate specificity observed in dipeptidyl-peptidase VI, YkfC and other NlpC/P60 domains in general. The first SH3b domain plays an important role in defining substrate specificity by contributing to the formation of the active site, such that only murein peptides with a free N-terminal alanine are allowed. A conserved tyrosine in the SH3b domain of the YkfC subfamily is correlated with the presence of a conserved acidic residue in the NlpC/P60 domain and both residues interact with the free amine group of the alanine. This structural feature allows the definition of a subfamily of NlpC/P60 enzymes with the same N-terminal substrate requirements, including a previously characterized cyanobacterial L-alanine-γ-D-glutamate endopeptidase that contains the two key components (an NlpC/P60 domain attached to an SH3b domain) for assembly of a YkfC-like active site

    Heavy metal and nitrogen concentrations in mosses are declining across Europe whilst some “hotspots” remain in 2010

    Get PDF
    In recent decades, naturally growing mosses have been used successfully as biomonitors of atmospheric deposition of heavy metals and nitrogen. Since 1990, the European moss survey has been repeated at five-yearly intervals. In 2010, the lowest concentrations of metals and nitrogen in mosses were generally found in northern Europe, whereas the highest concentrations were observed in (south-)eastern Europe for metals and the central belt for nitrogen. Averaged across Europe, since 1990, the median concentration in mosses has declined the most for lead (77%), followed by vanadium (55%), cadmium (51%), chromium (43%), zinc (34%), nickel (33%), iron (27%), arsenic (21%, since 1995), mercury (14%, since 1995) and copper (11%). Between 2005 and 2010, the decline ranged from 6% for copper to 36% for lead; for nitrogen the decline was 5%. Despite the Europe-wide decline, no changes or increases have been observed between 2005 and 2010 in some (regions of) countries

    The Sorcerer II Global Ocean Sampling Expedition: Expanding the Universe of Protein Families

    Get PDF
    Metagenomics projects based on shotgun sequencing of populations of micro-organisms yield insight into protein families. We used sequence similarity clustering to explore proteins with a comprehensive dataset consisting of sequences from available databases together with 6.12 million proteins predicted from an assembly of 7.7 million Global Ocean Sampling (GOS) sequences. The GOS dataset covers nearly all known prokaryotic protein families. A total of 3,995 medium- and large-sized clusters consisting of only GOS sequences are identified, out of which 1,700 have no detectable homology to known families. The GOS-only clusters contain a higher than expected proportion of sequences of viral origin, thus reflecting a poor sampling of viral diversity until now. Protein domain distributions in the GOS dataset and current protein databases show distinct biases. Several protein domains that were previously categorized as kingdom specific are shown to have GOS examples in other kingdoms. About 6,000 sequences (ORFans) from the literature that heretofore lacked similarity to known proteins have matches in the GOS data. The GOS dataset is also used to improve remote homology detection. Overall, besides nearly doubling the number of current proteins, the predicted GOS proteins also add a great deal of diversity to known protein families and shed light on their evolution. These observations are illustrated using several protein families, including phosphatases, proteases, ultraviolet-irradiation DNA damage repair enzymes, glutamine synthetase, and RuBisCO. The diversity added by GOS data has implications for choosing targets for experimental structure characterization as part of structural genomics efforts. Our analysis indicates that new families are being discovered at a rate that is linear or almost linear with the addition of new sequences, implying that we are still far from discovering all protein families in nature
    corecore