How Cal.com Rebuilt AppSec After Going Closed Source
How Cal.com consolidated noisy security tooling into one continuous, context-aware pull request security program with Gecko.
Gecko Security

If you’ve spent time with static analysis tools, you’ve probably noticed they excel at finding certain vulnerability classes while completely missing others. SQL injection, XSS, and path traversal get caught reliably. Authentication bypasses, broken access control, and privilege escalation slip through almost universally.
This pattern isn’t random. It isn’t something that can be fixed by tuning rules or adding more signatures. It reflects a fundamental architectural constraint in how most static analysis tools represent code.
Most SAST scanners, including many of the newer tools that have AI capabilities, are built on Abstract Syntax Trees (ASTs). ASTs provide a syntactically complete representation of source code, but they’re semantically empty in ways that matter for security analysis. Understanding this limitation helps explain why entire classes of vulnerabilities remain invisible to static analysis, and what would need to change to address them.
When a SAST tool parses code into an AST, it constructs a tree structure representing the syntax. Function declarations, variable assignments, method calls, control flow statements. For a simple function, the AST might look something like this:
<code class="hljs language-yaml"><span class="hljs-attr">FunctionDeclaration:</span> <span class="hljs-string">"get_order"</span>
<span class="hljs-string">├──</span> <span class="hljs-attr">Parameters:</span> [<span class="hljs-string">order_id</span>, <span class="hljs-string">current_user</span>]
<span class="hljs-string">├──</span> <span class="hljs-attr">Body:</span>
<span class="hljs-string">│</span> <span class="hljs-string">├──</span> <span class="hljs-attr">MethodCall:</span> <span class="hljs-string">"order_service.fetch"</span>
<span class="hljs-string">│</span> <span class="hljs-string">│</span> <span class="hljs-string">└──</span> <span class="hljs-attr">Arguments:</span> [<span class="hljs-string">order_id</span>]
<span class="hljs-string">│</span> <span class="hljs-string">└──</span> <span class="hljs-string">ReturnStatement</span>
</code>This representation captures the structure of the code within that file, but it doesn’t capture several things that matter for security analysis:
order_service.fetch actually defined?current_user ever used for authorization anywhere in the call chain?order_service.fetch, and do any of them include permission checks?These questions require information that ASTs don’t contain. An AST tells you the syntactic structure of code in a single file, but nothing about what that code actually does or how it connects to the rest of the system. For vulnerability classes that depend on understanding relationships between components, this limitation is fundamental.
Business logic vulnerabilities differ from injection vulnerabilities in an important way. Injection vulnerabilities are typically about something being present that shouldn’t be, like unsanitized input reaching a dangerous sink. Business logic vulnerabilities are often about something being absent that should be present. A missing authorization check, a validation that doesn’t happen, an assumption that doesn’t hold.
Take the authentication bypass we recently found in Cal.com. The vulnerability required chaining three separate bugs across different files:
No single file contained the vulnerability. Each function looked reasonable when examined in isolation. The bug existed only in the interaction between them, a three-step chain where each step’s flawed assumption enabled the next.
An AST-based scanner examining any of these files individually would see a function that queries the database, conditional logic that checks organization membership, and an upsert operation with proper error handling. All syntactically valid. All part of a critical authentication bypass that the scanner can’t detect.
If you’re familiar with SAST internals, taint analysis is the natural solution to this limitation. Serious tools don’t just parse ASTs; they track how data flows through the program to identify when untrusted input reaches dangerous operations.
Good taint analysis tools already have sophisticated symbol resolution. Tools like CodeQL construct a full semantic database with types, call graphs, and cross-file information, because you can’t do meaningful inter-procedural taint tracking otherwise. But better symbol resolution doesn’t fix the underlying problem, because taint analysis is designed to answer a specific question, and that question isn’t the one business logic vulnerabilities pose.
Taint analysis asks whether untrusted data reaches a dangerous operation without sanitization. This works well for injection vulnerabilities like SQL injection and XSS, which are fundamentally data-flow-to-dangerous-sink problems. But taint analysis can’t tell you whether the authorization logic is correct.
The Cal.com authentication bypass illustrates this. The taint trace showed request.email flowing through usernameCheckForSignup(), then through validateUsername(), then to prisma.user.upsert(). From a taint perspective this looks correct. Input comes from the request, passes through validation functions, and reaches the database.
The bug was inside the validation function:
;<code class="hljs language-typescript">
<span class="hljs-keyword">if</span>{' '}
(userIsAMemberOfAnOrg){' '}
{
<span class="hljs-comment">
// Skip validation entirely, leave available: true
</span>
}
</code>Taint analysis doesn’t evaluate whether conditional branches implement correct security policy. It tracks what flows through the program, not whether the logic is right. This is a fundamental limitation. Business logic vulnerabilities aren’t data-flow questions. They’re questions about correctness, about whether the code does what it should rather than just whether data flows where it shouldn’t.
| Question | Taint Analysis | Semantic Model |
|---|---|---|
| Does user input reach the database? | Yes | Yes |
| Is input sanitized before SQL execution? | Yes | Yes |
| Is there an authorization check in this path? | No | Yes |
| Is the authorization logic actually correct? | No | Yes |
| Does user context get dropped between layers? | No | Yes |
| Should this parameter be checked but isn’t? | No | Yes |
This is why adding an LLM on top of taint analysis doesn’t suddenly enable business logic detection. The taint trace for the Cal.com bug looks clean. There’s nothing for the LLM to flag because the underlying model doesn’t surface the actual problem.
The gap isn’t in parsing or tracing. It’s in reasoning about correctness. To find business logic vulnerabilities, you need a system that can:
order_service.get() in repo A calls handlers.get() in repo B.user is available at the entry point but never passed downstream.order.user_id is checked against current_user.id.if (userIsAMemberOfAnOrg) branch implements correct security policy.This is where the code representation matters. Feed an LLM incomplete or inaccurate code structure and it’s guessing. Give it compiler-accurate symbol resolution, type information, and full call chains, and it can actually reason about whether the code is correct.
In a small monolithic codebase, you might get away with AST-level analysis if your call chains are short and the entire codebase fits in context. But modern applications are increasingly multi-repo, with shared libraries, internal packages, and microservices in separate repositories. They’re polyglot, with Python services calling Go services calling JavaScript frontends. And they’re distributed, with API calls, message queues, and gRPC connections that don’t appear in any single AST.
Consider a typical authorization pattern across services:
<code class="hljs language-python"><span class="hljs-comment"># api-gateway/routes.py:</span>
<span class="hljs-meta">@app.get(<span class="hljs-params"><span class="hljs-string">"/orders/{order_id}"</span></span>)</span>
<span class="hljs-keyword">def</span> <span class="hljs-title function_">get_order</span>(<span class="hljs-params">order_id: <span class="hljs-built_in">str</span>, user: User = Depends(<span class="hljs-params">get_current_user</span>)</span>):
<span class="hljs-keyword">return</span> order_service.get(order_id) <span class="hljs-comment"># Does this check user owns order?</span>
</code><code class="hljs language-python"><span class="hljs-comment"># order-service/handlers.py (different repo):</span>
<span class="hljs-keyword">def</span> <span class="hljs-title function_">get</span>(<span class="hljs-params">order_id: <span class="hljs-built_in">str</span></span>):
<span class="hljs-keyword">return</span> repo.find_by_id(order_id) <span class="hljs-comment"># No user context here at all</span>
</code>The AST for each file is complete. But the vulnerability, that user authorization happens at the gateway but isn’t enforced at the service layer, is invisible to any tool that can’t connect these two files across repository boundaries.
AST-based tools can’t resolve that order_service.get maps to handlers.py:get(). They can’t trace that the user parameter is never passed downstream. They can’t determine that find_by_id returns data without ownership validation.
IDEs solved the cross-file resolution problem years ago. When you command-click on a function in VS Code, it doesn’t pattern-match the function name. It uses a language server to resolve the actual definition across files and dependencies.
Language servers build a semantic index of code. Every symbol, its type, where it’s defined, where it’s referenced, and how it connects to everything else. This is the same information compilers use, and it’s accurate, complete, and cross-file by design.
With a semantic index, you can answer questions like:
Semantic indexing solves cross-file analysis within a repository. For microservices, you need an additional layer to link services through their API contracts.
Most services define their interfaces somewhere, whether in OpenAPI specs, protobuf schemas, AsyncAPI definitions, or route decorators in the code itself. By parsing these contracts and mapping HTTP clients to the REST endpoints they call, message publishers to subscribers on those topics, and gRPC clients to service definitions, you can extend the semantic graph across service boundaries.
When you trace a call chain, it doesn’t stop at http.post(“/api/orders”). It continues into the service that handles that endpoint.
This is how you can detect vulnerabilities like:
AI coding assistants figured this out already. Copilot, Cursor, and similar tools use semantic indexing through IDE integrations or systems like GitHub’s stack graphs because AST-level understanding isn’t enough to write useful code. The same applies to finding vulnerabilities in it.
At Gecko, we built our scanner on language server indexing rather than AST parsing. The semantic index provides the ground truth of how code actually connects across files, repositories, and services.
On top of that foundation, we use LLMs to do what they’re suited for. Reasoning about developer intent, identifying security-relevant patterns, and generating targeted test cases. The LLM isn’t guessing about code structure. It’s reasoning about security using accurate information about how the code actually behaves.
The result is that we can find vulnerabilities like the Cal.com authentication bypass. Multi-step chains where each individual function looks correct, but the interaction between them creates an exploitable flaw.
If you’re dealing with business logic vulnerabilities that your current tools miss, or you’re trying to get visibility across a microservice architecture, you can try Gecko for free and see what it finds.

Jeevan “JJ” Jutla
Co-founder & CEO
JJ joined GCHQ as a teenager, working on security research and exploit development, and scored the highest mark on its reverse engineering and binary exploitation aptitude test ever recorded. The record still stands. He went on to lead security tool development for Binance’s eight-person red team and led the recovery of $500K stolen by North Korean state hackers, the largest recovery of stolen funds at the time.
The latest news, technologies, and resources from our team.
How Cal.com consolidated noisy security tooling into one continuous, context-aware pull request security program with Gecko.
Gecko Security
Authorization bypass in n8n’s dynamic-credentials OAuth endpoints allows any authenticated user to operate on another user’s OAuth credential by supplying its ID, enabling unauthorized OAuth rebinding and revocation.
Artemiy Malyshau
An IDOR vulnerability in n8n’s public variables API allows authenticated users to read project variables outside their authorized scope, exposing secrets across project boundaries.
Artemiy Malyshau
Learn API scanning for automated security testing. Find vulnerabilities from broken authentication to business logic flaws in your endpoints.
Artemiy Malyshau
A complete guide to automated pentest tools and best practices. Learn what works, what doesn’t, and how to implement continuous security testing.
Artemiy Malyshau
Compare the best AI-powered application security testing tools. Find which tools detect business logic flaws and broken access control.
Artemiy Malyshau
Occasional updates, new content, and insights. No spam; unsubscribe anytime.