STRING is a protein association network database that shows known and predicted relationships between proteins. It combines evidence from experimental assays and computational predictions to help researchers study protein interactions, functional modules, biological pathways, and protein complexes while evaluating the strength of supporting evidence.
The term STRING can be confusing because it may refer to a physical cord, a programming data type, a mathematical concept, or a biological resource. This article focuses on STRING as a protein association network database and explains how it helps researchers interpret relationships between proteins, evaluate evidence behind interactions, and decide when network results are useful.
STRING is a database that collects and analyzes protein association networks, showing known and predicted relationships between proteins. It does not simply list confirmed physical contacts; it combines multiple evidence sources, including experimental assays and computational predictions, to help researchers explore functional connections, protein complexes, and biological pathways.
STRING database explained: what it is and why it matters
STRING refers to a protein association network database used in bioinformatics to study how proteins are functionally connected. Researchers use it to investigate questions such as which proteins may participate in the same biological process, which groups of proteins form functional modules, and which relationships deserve further experimental study.
The word "STRING" has other meanings outside biology. A physical string is a thin cord or thread, while in programming a string is a sequence of characters stored as text data. Mathematics may also use the word in specialized contexts. When the topic involves proteins, genes, pathways, or bioinformatics, STRING refers to the biological database rather than these unrelated meanings.
Meaning of STRING by context
| If | Then |
|---|---|
| The search topic involves proteins, biology, or bioinformatics | Use STRING as the protein association database for understanding functional relationships between proteins. |
| The search topic involves textiles or physical materials | Use the general meaning of the word string rather than the biological database. |
| The search topic involves programming syntax or mathematics | Use the technical meaning of string in that field rather than the protein network resource. |
A protein association network represents proteins as connected points in a system. Each protein can be viewed as a node, while the connections between nodes represent evidence that proteins are related through shared functions, pathways, interactions, or biological context.

This approach matters because biology rarely depends on a single protein acting alone. Protein complexes and functional modules often involve groups of molecules working together, so network analysis can reveal relationships that are difficult to understand by examining isolated proteins.
How STRING builds protein association networks
STRING builds networks by collecting different types of biological evidence and integrating them into relationships between proteins. Some links come from direct experimental observations, while others are predicted through computational approaches that identify possible functional connections.
A computational prediction does not mean a protein relationship has been proven. Instead, it indicates that available biological information suggests a connection worth investigating.
In practical bioinformatics analysis, the easy-to-miss step is separating a network connection from a confirmed biological mechanism. A line between two proteins on a network diagram is evidence to evaluate, not automatic proof that the proteins physically interact in every biological condition.
STRING can incorporate computational approaches that use patterns such as genomic relationships, sequence information, and known biological associations. These methods help expand networks beyond the interactions that have been directly observed in laboratory settings.

The general workflow is:
- Collect evidence from experimental and computational sources.
- Evaluate how strongly each evidence type supports a protein association.
- Combine evidence into interaction records displayed in the network.
- Allow researchers to explore clusters, hubs, and functional relationships.
A protein complex is a group of proteins that work together as a biological unit. Functional modules are groups of related proteins that participate in connected biological activities. STRING networks can help identify these patterns by showing groups of proteins that repeatedly appear as connected communities.
Experimental and predicted evidence behind STRING interactions
STRING interactions are easier to interpret when evidence types are separated. A researcher should first ask what kind of support exists behind a connection rather than treating every edge in a network as equal.
Evidence behind STRING interactions
| Evidence type | Data source | Typical role | Interpretation |
|---|---|---|---|
| Experimental assay evidence | Laboratory-based interaction experiments | Provides observed biological support for an association | Higher-confidence starting point for generating hypotheses |
| Computational prediction evidence | Sequence, genomic, or computational relationship models | Suggests possible protein relationships for investigation | A lead that requires independent confirmation before being treated as a conclusion |
| Combined interaction evidence | Multiple evidence channels integrated into one network record | Ranks and displays relationships using accumulated support | Useful for prioritizing research questions while keeping evidence strength in context |

Experimental assays provide observations collected from laboratory methods. These results can support the idea that proteins are connected under particular testing conditions, but interpretation still depends on the biological system and experimental design.
Computationally predicted interactions fill gaps where direct measurements are unavailable. They are valuable for generating hypotheses, finding possible pathways, and selecting candidates for follow-up work. However, a predicted interaction should not be treated as confirmed fact without additional validation.
From an editorial review of biological database explanations, the recurring failure mode is presenting a network as if it were a simple map of proven interactions. The more useful question is: what evidence created this connection, and is that evidence appropriate for the research decision being made?
Understanding STRING confidence scores in practice
STRING confidence scores are designed to communicate how much supporting evidence exists for a protein association. The practical meaning is not that a score turns an association into biological truth; rather, it helps users compare stronger and weaker evidence-supported relationships within a network.
A higher confidence relationship generally indicates that more supporting evidence has been accumulated. A lower confidence relationship may still be useful for exploration, but it should be treated as a hypothesis generator rather than a final conclusion.
A safer way to interpret an interaction is to check three things:
- Identify whether the evidence is experimental, predicted, or combined from multiple sources.
- Compare the confidence level with the purpose of the analysis.
- Require additional biological validation before making strong claims from uncertain predicted relationships.
For example, a researcher searching for possible proteins involved in a disease pathway may use lower-confidence connections to discover candidates for investigation. The same evidence may be insufficient if the goal is to claim that a specific protein directly controls a biological process.

The important distinction is between prioritization and proof. STRING can help decide which relationships deserve attention, but the database does not replace experimental confirmation.
Using STRING networks for biological research questions
STRING supports several research applications because protein networks can reveal relationships that are difficult to see from individual protein records. Common uses include studying biological pathways, exploring protein complexes, investigating disease-related mechanisms, and identifying possible drug target candidates.
For disease diagnosis research, STRING can support the investigation of whether groups of proteins are connected to biological processes associated with a disease. The network may help researchers identify relevant proteins or pathways to examine further, but it does not by itself provide a clinical diagnosis.
For drug target identification research, STRING can help researchers examine whether a protein is connected to other proteins involved in a biological process of interest. A highly connected protein may become a candidate for additional study, but network importance alone does not prove that changing that protein will produce a therapeutic effect.
Reading STRING visualization outputs
A STRING visualization represents relationships as a network diagram. Nodes usually represent proteins, while connecting lines represent associations supported by available evidence.
Different visualization choices answer different questions. A researcher looking for major connected groups may focus on clusters of nodes, while someone studying a single protein may examine its surrounding connections and evidence sources.
A hub protein is a protein that has an unusually large number of connections within a protein association network. Researchers often examine hub proteins because highly connected nodes may represent important biological functions or candidates for further study.
A hub status alone is not enough to conclude that a protein is a disease marker or drug target. It is a signal for investigation, not a final biological decision.
Applying STRING to functional analysis
Functional analysis uses network relationships to explore whether connected proteins participate in shared biological processes. A cluster of related proteins may suggest a pathway, protein complex, or functional module that deserves closer examination.
A practical interpretation workflow is:
- Start with a protein or group of proteins connected to the research question.
- Review the network connections and the evidence supporting important links.
- Look for clusters or modules that suggest shared biological functions.
- Confirm important findings with appropriate biological experiments or complementary resources.

This approach is useful when the research question involves functional relationships between proteins. It is less suitable when the question depends only on information outside protein association networks, such as measurements that require another specialized data source.
Limitations of STRING predicted interactions
Predicted interactions are one of STRING’s most useful features and one of the easiest areas to misunderstand. They expand biological networks beyond known experimental observations, but they also introduce uncertainty.
The main failure mode is treating a computational prediction as a confirmed interaction. The warning sign is when a network connection is cited as biological fact without checking the underlying evidence type or confidence information.
Predictions can be affected by limitations in the underlying data, assumptions in computational models, or differences between biological conditions. A relationship predicted from general biological patterns may not occur in a specific cell type, organism, or experimental situation.
Use an evidence-aware workflow:
- Separate predicted interactions from experimentally supported interactions before interpreting results.
- Review the strength and type of evidence behind important network links.
- Seek independent confirmation before using low-confidence predictions as research conclusions.
This does not make predicted interactions unusable. They are valuable for generating ideas, prioritizing experiments, and exploring possible biological relationships, as long as their uncertainty remains visible.
When STRING is the right research resource
STRING is a strong fit when the research question involves protein-level functional relationships. It is especially useful when researchers need to explore how proteins connect, which groups of proteins may work together, or which relationships may deserve further investigation.
Before using STRING, match the database to the question rather than starting with the tool itself.
- Confirm that the research question involves protein-level functional relationships.
- Verify that the goal is not limited to gene expression measurements alone.
- Check whether network-based evidence can answer the biological question being studied.
- Identify any complementary resources needed for data types outside STRING’s scope.
A research workflow may combine STRING with other databases, laboratory methods, or analytical tools depending on the question. STRING is most useful when network context adds information that an isolated protein list cannot provide.
When you are choosing a resource today, enter one protein or protein group related to your question into STRING and inspect the evidence types behind the first connections you see; this will show whether the database provides the kind of network evidence your project actually needs before you build conclusions from it.
FAQ
What does STRING stand for in biology?
In biology, STRING refers to a protein association network database used in bioinformatics to study functional relationships between proteins. It is not referring to the general meaning of a physical string or a programming string. The database helps researchers explore known and predicted protein associations using different evidence sources.
Is STRING a database or a software tool?
STRING is a protein association network database that collects and analyzes relationships between proteins. Researchers use its network information and visualization features to explore interactions, functional modules, protein complexes, and evidence behind protein associations.
Can STRING show only experimentally confirmed protein interactions?
No. STRING includes both experimentally supported associations and computationally predicted relationships. A connection in the network represents evidence of an association, but it does not automatically prove a direct physical interaction. Researchers should review evidence types and confidence before drawing conclusions.
Is STRING free to use for researchers?
The provided information describes STRING as a database used by researchers for exploring protein association networks, evidence types, and biological relationships. It does not provide specific details about access terms or usage conditions, so availability should be checked through the resource itself.

