Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Yes, sensitive corporate data has been found in prompts and files submitted to AI tools. But available figures do not show what share of developers do this: one vendor-monitored sample is not a representative survey, and statistics about credentials leaked to public code repositories measure a different problem.
To assess the risk, separate three paths: a developer submits information directly, an AI assistant or agent can access workspace material, or a credential is committed to a repository. Each needs different controls.
What does the evidence say about developers sharing secrets with AI?
Axios reported on July 31, 2025, that Harmonic Security analyzed one million prompts and 20,000 files submitted to 300 AI tools and AI-enabled SaaS applications between April and June 2025. In that sample, more than 4% of prompts and more than 20% of uploaded files contained sensitive corporate data; code was the most common sensitive-data type reported in prompts. Axios’s report says the sample came from organizations using Harmonic’s tools. It is evidence that exposure happens, not a representative estimate of all organizations or developers.
Free tools Windows power users keep installed
One-click scans. No signup required.
No representative developer-specific prevalence figure is established here. In particular, repository leak counts cannot fill that gap: they track credentials found in source-control repositories, not what people type or upload to LLMs.
#1 Best Overall
Keep repository statistics in their lane
GitHub reported that more than 39 million secrets were leaked across GitHub in 2024. That is a repository-leak figure, not a count of secrets pasted into AI prompts. GitHub’s 2025 report describes that repository finding. Separately, GitHub said more than one million leaked secrets were detected on public repositories in the first eight weeks of 2024; that, too, is a repository statistic. GitHub’s 2024 post covers that finding.
How can code or credentials reach an AI system?
| Exposure path | What may be exposed | What to control |
|---|---|---|
| Direct prompt or file submission | Text, code, logs, configuration, or an uploaded document that a developer intentionally includes. | Approved tools, acceptable data classes, and habits for removing credentials and unnecessary proprietary context before submission. |
| Assistant or agent workspace access | Code and other material an assistant or agent can access through its workspace, repository permissions, or connected tools, even if the user does not paste each item into a prompt. | Limit access to the repositories, files, and tools needed for the task; assess what context the service can use and transmit. |
| Repository credential leak | A secret committed to source control, where it may be detected by scanning or blocked by push protection. | Secret scanning, push protection, alert ownership, and a response process for exposed credentials. |
These paths can overlap, but a safeguard for one does not automatically cover the others. For example, repository secret scanning is useful for credentials committed to code; it does not filter every prompt sent to an external AI service.
Rank #2
- Used Book in Good Condition
What should teams verify about an AI tool’s data handling?
Do not assume that all AI products—or every plan and configuration within one product—handle interactions the same way. Review the actual service, subscription, organization settings, and any selected provider. GitHub’s Copilot product information says interaction-data treatment depends on plan and notes that data from individual subscribers may be used to train and improve models. GitHub’s responsible-use documentation also says that, in a bring-your-own-key setup, prompts and responses are transmitted to the selected provider and may be subject to that provider’s retention and privacy policies. These statements are not a rule for every AI service or account; verify the terms that apply to yours.
- What prompts, files, repository content, or workspace context is sent or made accessible?
- Under this plan, can interaction data be used for model training or improvement?
- What retention and deletion terms apply, including at a bring-your-own-key provider?
- Which administrators can approve tools, manage user access, and set agent permissions?
- Can the organization detect or block secrets committed to repositories, and who receives the alerts?
Why do AI agents need additional safeguards?
An agent can act on information it reads, not just the words a developer deliberately types. NIST’s Center for AI Standards and Innovation describes agent hijacking as indirect prompt injection: malicious instructions are placed in data an agent may ingest, taking advantage of the lack of a clear separation between trusted instructions and untrusted content. In a January 17, 2025 evaluation, CAISI added tests for remote code execution, database exfiltration, and automated phishing, and said it was frequently able to induce agents to follow malicious instructions across those new risk areas. This describes that evaluation; it does not establish that every current agent is vulnerable in the same way. See NIST’s evaluation account.
Rank #3
GitHub’s documentation for its cloud agent similarly warns that an agent with access to code and sensitive information could leak it accidentally or in response to malicious user input. GitHub’s documented risks and mitigations are specific to that product context, but the practical lesson is broader: grant an agent only the access needed for its task, and test how it handles untrusted content in workflows where it can take consequential actions.
What should an organization do to reduce exposure?
- Define approved tools and data classes. Specify which AI services employees may use and what kinds of company information may be submitted. Make the boundary clear for credentials, customer data, source code, and other confidential material.
- Reduce what developers submit. Train developers to remove credentials and unnecessary proprietary context before sending a prompt or file. Where possible, use redacted examples or synthetic data instead of live secrets.
- Verify terms and configuration. Check the current plan-specific and provider-specific rules for training or improvement, retention, deletion, and organization controls. Recheck them when the service, plan, or setup changes.
- Constrain agent permissions. Limit repository, file, and tool access to what the task requires. Review whether an agent can read sensitive material or take actions that could expose or misuse it.
- Test realistic agent workflows. Include untrusted instructions embedded in content the agent may read, and check whether it reveals data or performs actions beyond the intended task.
- Use repository protections as one layer. Enable secret scanning and push protection where available, route alerts to an owner, and have a process for responding to discovered credentials. These controls address repository exposure, not every prompt submission.
For broader confidentiality incident handling, NIST SP 1800-28 addresses identifying and protecting data, while SP 1800-29 addresses detecting, responding to, and recovering from confidentiality attacks. They are general guidance, not LLM-specific standards: SP 1800-28 and SP 1800-29.
Rank #4
Which NIST guidance applies to secure AI development?
NIST SP 800-218A, finalized July 26, 2024, supplements the Secure Software Development Framework with practices for AI model development across the software development lifecycle. NIST says it is intended for producers of AI models, producers of AI systems that use models, and acquirers of those systems. It can help organizations frame secure-development responsibilities, but it is not evidence of how often employees submit secrets to chatbots. Read the SP 800-218A profile.
NIST’s Control Overlays for Securing AI Systems project identifies proposed use cases that include adapting and using an LLM assistant, using single- or multi-agent systems, and security controls for AI developers. The project page reported a concept paper available for comment on August 14, 2025; check the project page for its current status rather than treating the overlays as final requirements.

