Structure-aware Tm and OGT prediction with a calibrated interval
StableProt couples SaProt 3Di conformational tokens with disjoint heads for melting temperature (Tm) and optimal growth temperature (OGT). It returns a calibrated predictive interval rather than a point estimate. Training and every evaluation split are separated by a bidirectional homology audit (<30% sequence identity), with mesophilic downsampling on the OGT set.
StableProt is a structure-aware model for protein melting temperature (Tm) and organismal optimal growth temperature (OGT). The Tm and OGT heads are disjoint: the OGT prior enters the Tm path only. Predictions average a 5-seed ensemble and include a scaled predictive interval.
Training corpus comprises 28,739 clean Tm records (curated from 29,300 after permanently purging 561 sequences sharing ≥30% identity with evaluation constructs) and 940,000 curated OGT records with mesophilic downsampling. Held-out evaluation benchmarks include ProThermDB (n = 3,340), FireProtDB (n = 322), and BRENDA / BacDive OOD test sets under strict bidirectional homology decontamination (<30% sequence identity).
On bin-balanced OGT scoring across the thermal spectrum, StableProt maintains robust accuracy across all environmental regimes without degradation outside the mesophilic band. On FireProtDB single-point mutations (n = 3,649 physical experimental mutations), ΔTm MAE is 5.05 °C with 56.8% directional classification accuracy (ROC AUC = 0.585). As a sequence-level stability model not trained on mutational assay differentials, its predicted interval widens appropriately to reflect single-mutation uncertainty.
Loop highlighting uses Chou–Fasman propensities, then re-scores Tm for edited sequences. It is a sequence-side aid, not a wet-lab-validated design method.
Computational Biology and Bioinformatics Laboratory, iBRIC–Institute of Life Sciences (ILS), Bhubaneswar, Odisha, India.
Contact: bibhu.prasad@ils.res.in and anshuman@ils.res.in.
This server is for academic and non-commercial research. Commercial use needs written permission from the authors. Predictions are provided as-is, with no warranty. They are not a clinical, diagnostic, or wet-lab result. The Design tab is a sequence-side hypothesis aid.
The served model is a 5-seed ensemble. The half-width shown is ±1σ (about 68% coverage in distribution), not a 95% confidence interval.
Sequences are sent to this server only to run inference. The application does not write submitted sequences to disk. The host may keep ordinary web-server access logs (time, IP, path) that do not include the sequence body.
The page loads fonts from Google Fonts, which is a third-party request from your browser.
No journal citation yet. Until then, cite the working manuscript and the repository: