Incident Readiness You and Your Leaders Will Actually Trust
Shachar Hirshberg (CEO and Co-founder · Artemis), Hadar Waldman (Security Engineer · Artemis)
BSidesSF 2026 · Day 1 · AMC Theatre 13
Overview
In an era where production environments are increasingly complex and dynamic, ensuring robust incident readiness and maintaining visibility into operational realities has become a critical challenge for cybersecurity teams. Shachar Hirshberg and Hadar Waldman from Artemis tackled this pressing issue in their BSides SF talk, "Incident Readiness You and Your Leaders Will Actually Trust." Their presentation highlights a significant disconnect between an organization's security policies and the actual state of its production systems, proposing an innovative, data-driven solution leveraging Large Language Models (LLMs) to bridge this gap.
Key moments
- 0:00 Speakers' introduction and talk overview
- 2:30 The challenge of gaining visibility in production
- 5:37 Introducing the environmental intelligence report solution
- 6:07 Practical steps to create the intelligence report
- 7:00 Analyzing identity logs for privileged actions
- 8:00 Identifying unusual authentications via geographic data
Incident Readiness You and Your Leaders Will Actually Trust
Speakers: Shachar Hirshberg, CEO & Co-founder, Artemis; Hadar Waldman, Security Engineer, Artemis
Conference: BSides SF
YouTube: https://www.youtube.com/watch?v=Cz5e9Oc_pQ8
Overview
In an era where production environments are increasingly complex and dynamic, ensuring robust incident readiness and maintaining visibility into operational realities has become a critical challenge for cybersecurity teams. Shachar Hirshberg and Hadar Waldman from Artemis tackled this pressing issue in their BSides SF talk, "Incident Readiness You and Your Leaders Will Actually Trust." Their presentation highlights a significant disconnect between an organization's security policies and the actual state of its production systems, proposing an innovative, data-driven solution leveraging Large Language Models (LLMs) to bridge this gap.
The core of their approach revolves around generating what they term an Environmental Intelligence Report. This report is derived from the systematic analysis of an organization's existing log data using LLMs, providing deep, actionable insights that traditional security tools often miss. The speakers emphasize that by regularly producing these intelligence reports, security teams can not only uncover hidden risks and inefficiencies but also build trust with leadership through consistent, impact-driven communication, ultimately fostering a more secure and cost-effective operational posture.
This talk is particularly relevant for security professionals struggling with the sheer volume and complexity of telemetry data, those seeking to move beyond static compliance checks, and leaders aiming to gain a clearer, more trustworthy understanding of their organization's real-world security landscape. As production environments move towards increased automation and agent-driven operations, the ability to derive meaningful intelligence from raw data will become even more indispensable, making the methodology presented by Hirshberg and Waldman a timely and powerful paradigm shift in incident readiness.
Background
▶ Watch: Speakers' introduction and talk overview (0:00)
The pervasive challenge of achieving comprehensive visibility into production environments stems from their inherent complexity. Modern infrastructures, especially those leveraging cloud services, are vast, distributed, and constantly evolving. Organizations frequently deploy a myriad of security tools, including Cloud Security Posture Management (CSPM) systems and Security Information and Event Management (SIEM) platforms, to establish guardrails and monitor activity. However, as the speakers pointed out, a significant portion of security professionals (over half in their informal poll) are not confident that these tools provide a complete and accurate picture of what's truly happening in production.
The speakers articulated that while static analysis tools can confirm if a policy is configured, they often fail to validate if that policy is enforced in practice. For instance, a CSPM might report that Multi-Factor Authentication (MFA) is enabled for a given role, but it won't necessarily confirm if users are consistently utilizing MFA, or if specific legacy rules allow exceptions that undermine the overall security posture. This creates a critical "gap between the static configuration and the policy to reality in production" – a blind spot where misconfigurations, shadow IT, and unauthorized activities can thrive undetected.
Traditional SIEMs, while capable of ingesting vast amounts of log data, often struggle to connect disparate data points, reason through complex scenarios at scale, or synthesize actionable intelligence without extensive manual effort from highly skilled analysts. The sheer volume of data makes it challenging for humans to identify subtle anomalies or correlate events across different security domains (identity, network, endpoint, cloud). This problem is exacerbated by the anticipated future state where "most of activity in production will be operated largely by agents," further obscuring direct human oversight and demanding advanced, automated intelligence capabilities. The need for a solution that can not only ingest but also intelligently interpret and contextualize this deluge of operational telemetry is therefore paramount for effective incident readiness.
Key Findings
▶ Watch: Introducing the environmental intelligence report solution (5:37)
The central finding presented by Shachar Hirshberg and Hadar Waldman is the existence of a pervasive and often surprising "gap between the static configuration and the policy to reality in production." They consistently observe that organizations, particularly at enterprise scale, are frequently astonished by the discrepancies uncovered in their live environments compared to their documented policies or static security assessments. This revelation underpins the necessity of their Environmental Intelligence Report methodology.
Their research and practical application of this methodology have led to several key discoveries:
- Ubiquitous Hidden Risks: Across every organization Artemis has engaged with, they've found a "ton of stuff across identity, cloud, network, endpoint that nobody knew about." These hidden findings range from dormant privileged accounts and misconfigured IAM (Identity and Access Management) policies to unauthorized application usage and overlooked attack surfaces from deprecated services.
- LLMs as Transformative Analytical Tools: The speakers emphasize that LLMs are not merely a scheduling tool for reports but offer a transformative capability to "connect the dots across multitude of data points when they're given the right context." Unlike humans who struggle to analyze "hundreds of pages and hundreds of thousands of rows" at scale, LLMs can reason iteratively and in an agentic manner, correlating identity data with endpoint and network information to build a comprehensive picture of user activity and legitimacy.
- Significant Cost Savings Potential: Beyond security posture improvement, the methodology frequently uncovers opportunities for substantial cost savings. They reported that "six customers so far, each and every one of them were able to find things that saved them multi-million dollars in their AWS spend" by identifying over-provisioned resources, misconfigured IAM policies leading to excessive CloudTrail and S3 costs, and other inefficiencies.
- The Value of Consistency and Iteration: The power of the Environmental Intelligence Report compounds over time. By maintaining consistent reporting (e.g., weekly or monthly cadence) and tracking progress, organizations can visibly demonstrate improvements, build trust with leadership, and justify further investment in security programs. This iterative approach allows for continuous refinement and a proactive stance against evolving threats and misconfigurations.
- Actionable Insights from Existing Data: A crucial finding is that organizations don't necessarily need to invest in entirely new data collection mechanisms. The intelligence can be derived from existing log data sources – whether in S3, Splunk, or other platforms – when combined with intelligent LLM analysis. This makes the approach highly accessible and immediately impactful using current infrastructure.
Technical Deep Dive
▶ Watch: Practical steps to create the intelligence report (6:07)
The core technical innovation presented is the Environmental Intelligence Report, a process designed to systematically extract actionable security insights from an organization's existing log data using Large Language Models (LLMs). The methodology is iterative and consists of three main steps: query logs, LLM analysis, and report generation, repeated over time to track changes and improvements.
1. Querying Logs:
The first step involves querying logs from various sources. The speakers stress that organizations can leverage their existing infrastructure.
- Log Storage: Common log storage solutions include Amazon S3 for cloud-native environments or traditional SIEMs like Splunk.
- Querying Tools: For S3, Athena can be used for direct SQL-like queries. For Splunk, the Splunk API combined with SPL (Splunk Processing Language) allows programmatic access to log data. The key is the ability to run these queries via an API to enable automation.
- Query Management: Queries and their outputs need to be saved and managed. The speakers mentioned using Jupyter notebooks or simple scripts for this purpose, allowing for repeatable execution.
2. LLM Analysis and Report Generation:
Once the raw log data (or aggregated query results) is obtained, it is fed into an LLM for analysis and report generation.
- AI Coding Buddy: Tools like Claude (from Anthropic) or Codex (or similar LLMs) are used to analyze the query outputs.
- Report Format: Markdown is recommended as an effective format for LLMs to generate, which can then be exported to PDF or DOC.
What to Look For (Identity Logs Example):
Hadar Waldman provided a detailed breakdown of specific categories and questions to guide the analysis, using identity logs as a prime example:
- Privileged Actions: Instead of just listing administrators, the focus is on identifying actions that require high privileges. This can reveal:
- Dormant Admins: Accounts with high privileges that are never used.
- Unexpected Admin Behavior: Admins performing actions outside their typical scope, from unusual geographic locations, or using unexpected user agents.
- Service Account Misuse: An example was given of a service account named "read only" performing write actions, or an administrator using a service account from their home IP to bypass their own privileges.
- Geographic Authentications: Analyzing the origin of authentications helps identify:
- Top/Least Used Locations: Highlighting expected vs. unexpected login locations.
- VPN Usage: Identifying legitimate VPN use vs. potentially unauthorized logins from unusual countries.
- Application Usage: Understanding which applications employees are authenticating to via identity providers like Okta or Entra.
- MFA Usage: This area often reveals significant gaps between policy and reality:
- Deprecated Policies: Finding old rules allowing password-only authentication that were thought to be removed.
- Weak MFA Factors: Discovering that non-fishing-resistant MFA factors are still in use, contrary to security expectations.
- Failed and Denied Actions: These logs can surface:
- Deprecated Services: Credentials or users attempting to authenticate to services thought to be removed, indicating lingering attack surface.
- Credential Stuffing/Brute Force: Patterns of failed logins.
LLM Integration – Using "Skills" for Consistency:
A critical aspect of making LLM-driven reports reliable is ensuring consistency, which is achieved through "skills" or pre-shaped instructions given to the LLM agent.
- Skills Definition: A "skill" is a set of instructions placed in a specific file, invoked by a slash command (e.g.,
/generate report). This ensures the same analysis is performed repeatedly. - Tips for Consistent Reports:
- Templates: Use predefined templates for report structure, such as tables with specific columns for findings. This prevents the LLM from reinventing the format each time.
- Sub-agents: For large reports (e.g., 60 pages), break down the task using sub-agents. This allows the LLM to process smaller chunks of data, maintaining higher quality and consistency throughout the entire report, not just the initial sections.
- Validation Step: Incorporate a validation step within the skill where the LLM (or another agent) checks the quality of each section. If a section is too short or lacks detail, the relevant sub-agent can be re-triggered to improve it.
- Clear Writing Guidelines: Provide explicit instructions on the desired language, tone, and style for the report to ensure consistent output.
Effective Prompting Techniques:
To mitigate the risk of LLMs generating confident but incorrect "assumptions" (or "lies," as Hadar jokingly put it), specific prompting techniques are crucial:
- Require Citations: Force the LLM to reference the specific data points that support its statements. This improves reviewability and ensures data-backed conclusions.
- Enforce Hedging: Instruct the LLM to use softer, less definitive language (e.g., "may indicate," "suggests," "appears to be") when making inferences, separating fact from assumption.
- Separate Fact from Inference: Clearly label parts of the report as factual (e.g., "account X logged in Y times") versus inferences drawn from those facts.
- Identify Missing Information: If the LLM lacks a key piece of information, it should explicitly state this rather than attempting to fill the gap with fabricated data.
- Challenge Claims: For potentially malicious findings, prompt the LLM to also generate a benign explanation and then compare the likelihood of both scenarios.
Scaling Data Processing:
The speakers acknowledged the challenge of feeding massive volumes of data directly to LLMs. Artemis processes "over 10 petabytes per day" and "billions of events every hour." To manage this, they employ multiple layers to synthesize and filter data before it reaches the LLMs, ensuring that agents receive only the most actionable and relevant information, preventing "bankrupt[cy]" from excessive LLM costs. They also mentioned using LLMs as judges and employing "multiple layers of agentic evaluation," sometimes up to "seven even like five steps of agents checking each other," to increase efficacy, albeit at a higher computational cost.
Demo / Proof of Concept
▶ Watch: Analyzing identity logs for privileged actions (7:00)
While the talk did not feature a live, interactive demonstration of the Environmental Intelligence Report system, the speakers presented a compelling conceptual proof of concept through their recounted experiences and the methodology itself. They detailed how Artemis, in its early stages and with its customers, has successfully implemented this approach to uncover significant security issues and cost savings.
The "proof of concept" lies in the consistent findings they reported across diverse client environments: "every organization that we go in, we find a ton of stuff across identity, cloud, network, endpoint that nobody knew about." This empirical evidence from real-world deployments serves as validation for their methodology. They illustrated with concrete examples, such as the discovery of deprecated password-only authentication rules, service accounts performing unauthorized write actions, and multi-million dollar savings in cloud spend due to misconfigured IAM policies.
The talk emphasized that these results were achievable by leveraging existing log data and readily available tools, suggesting that organizations can replicate this success "at home" with their current infrastructure. The detailed explanation of query types, LLM prompting, and agentic workflows effectively outlines a repeatable process that has been proven to yield valuable security and operational insights.
Defensive Implications
▶ Watch: Identifying unusual authentications via geographic data (8:00)
The Environmental Intelligence Report methodology offers profound defensive implications, empowering security teams to move beyond reactive incident response and static compliance checks towards a proactive, data-driven security posture. The speakers outlined specific strategies for leveraging these insights to drive action within an organization:
- For Leadership and Management – Lead with Impact:
- Risk Quantification: Instead of presenting raw numbers (e.g., "300 over-privileged roles"), frame findings in terms of business impact and risk to the organization. This translates technical jargon into language that resonates with decision-makers.
- Consistency and Progress: Establish a regular cadence (weekly or monthly) for sharing these reports. This allows for showing "wins and the gaps this week and the plan for next week," demonstrating tangible progress in reducing risk and building trust. Consistent, visible improvement encourages further investment in security programs.
- For Technical Teams (e.g., Cloud Security) – Provide Technical Evidence:
- Actionable Detail: When collaborating with teams responsible for remediation (e.g., cloud security teams for AWS issues), provide precise technical evidence. This includes linking directly to "the exact logs that show that the activity is actually the ones that report says is happening."
- Collaborative Resolution: Track the resolution of identified improvements together with the technical teams. This fosters a shared responsibility model and ensures fixes are implemented effectively.
- Cost Savings as a Security Driver:
- Quantify Over-provisioning: The reports can identify unused or over-provisioned resources, particularly common in cloud environments. By analyzing real telemetry over extended periods (multiple months), teams can prove that resources (like backup roles that haven't triggered) are truly unutilized, leading to significant cost reductions in cloud spend, CloudTrail costs, S3 storage, and even SIEM ingestion costs. This financial benefit can be a powerful motivator for security initiatives.
- Proactive Threat Hunting and Vulnerability Management:
- Identify Dormant Attack Surface: Uncover deprecated services, lingering credentials, or misconfigured accounts that represent unnecessary attack surfaces.
- Detect Policy Violations in Real-time: Continuously monitor for deviations from security policies, such as non-fishing-resistant MFA factors or password-only authentication where MFA is mandated.
- Uncover Shadow IT and Insider Threats: Identify unauthorized application usage, suspicious geographic logins, or unusual service account activity (e.g., an administrator using a service account from their home network to gain unauthorized privileges), and "unauthorized AI usage in organization."
- Continuous Improvement Program:
- The methodology fosters a culture of continuous improvement by systematically identifying, reporting, and remediating security gaps. It helps organizations transition from a static, audit-driven security model to a dynamic, intelligence-led approach that constantly refines its posture based on real-world operational data.
By implementing these defensive strategies, organizations can not only enhance their incident readiness but also build a security program that is transparent, impactful, and trusted by all stakeholders.
Key Takeaways
- The gap between static security policies/configurations and the actual reality in production environments is widespread and often surprising, necessitating dynamic, data-driven validation.
- Organizations can leverage existing log data (e.g., from S3, Splunk) and readily available LLMs (like Claude or Codex) to generate powerful Environmental Intelligence Reports without requiring entirely new data collection infrastructure.
- LLMs, when properly structured with skills, sub-agents, and effective prompting techniques (e.g., requiring citations, hedging, separating fact from inference), can consistently and accurately synthesize actionable insights from vast and complex datasets.
- To drive action and build trust, security teams should frame findings in terms of business impact and consistently demonstrate progress through regular reports, providing technical evidence for remediation teams.
- Beyond security posture improvement, this methodology frequently uncovers significant cost-saving opportunities by identifying and quantifying over-provisioned resources and inefficient configurations in cloud environments.
- Despite the power of LLMs, human oversight remains critical; thorough review of LLM-generated reports is essential to ensure accuracy and maintain accountability.
About the Speaker(s)
Shachar Hirshberg is the CEO and co-founder of Artemis. With over 15 years of experience in technology and cybersecurity, Shachar has a distinguished background that includes significant contributions to major industry platforms. Notably, he spent multiple years at AWS, where he was instrumental in building Amazon GuardDuty, a continuous threat detection service. Prior to that, he also played a key role in developing Demisto and Source Space, both of which are prominent solutions in the cybersecurity automation space. At Artemis, he is focused on helping companies enhance their security posture by augmenting or replacing traditional SIEM solutions.
Hadar Waldman is a Security Engineer at Artemis. She brings 12 years of hands-on experience in various facets of detection and response, having served as a SOC analyst, detection engineer, SOC manager, and incident responder. Her diverse background gives her a deep understanding of the practical challenges and needs of security operations teams. Beyond her professional expertise, Hadar is passionate about data analytics and describes herself as a "big musicals geek," a personal detail she humorously shared during the talk.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
A competent BSides-level case study on using LLMs to surface the policy-vs-reality gap in production environments. The methodology is practical and the speakers have genuine operational credibility, but this is applied tooling guidance dressed up as novel research — the core insight (your CSPM lies to you, go look at actual logs) predates LLMs by a decade.
Heather Calloway (CISO) — SOLID
A practical, well-grounded talk on using LLMs to surface the gap between configured policy and operational reality — a real problem with real consequences. The methodology is accessible and the cost-savings angle is a smart bridge to leadership, but it stays at the practitioner level and never quite reaches the governance or accountability questions that make this problem institutional rather than technical.