Bilkent University
Department of Computer Engineering
M.S.THESIS PRESENTATION
SWE FIREWALL: PRIVACY-AWARE PROGRAM REPAIR WITH LARGE LANGUAGE MODEL AGENTS
Giray Akyol
Master Student
(Supervisor: Assoc.Prof.Eray Tüzün )
Computer Engineering Department
Bilkent University
Abstract: Cloud-hosted LLM program-repair systems can expose proprietary source code through both one-shot prompts and iterative agentic interactions. Objectives: We evaluate whether masking, source-level obfuscation, intent inference, and dynamic access control can reduce disclosure while retaining repair effectiveness. Methods: We evaluate eight one-shot context configurations on 213 Defects4J bugs using GPT-5.2, Llama Scout, and Gemini 2.5 Flash-Lite. We evaluate Explorer Subagent and File Budget controls on 500 SWE-bench Verified tasks using GPT-5.2, Grok 4.3, and Grok 4.1 Fast. Repair is measured by plausible@1 for one-shot configurations and resolved-task rate for agentic configurations. Reconstruction is measured using average token log probabilities and CodeBLEU; agentic disclosure is additionally measured using file cost, unlock cost, and collateral compression. Results: One-shot FullContext achieves 23.01% plausible@1, compared with 20.03% for IntentContext, 22.22% for ObfFullContext, and 17.84% for ObfIntentContext. IntentContext changes the average reconstruction log probability from −0.0212 to −1.0153 and CodeBLEU from 0.3037 to 0.2716; ObfIntentContext obtains CodeBLEU 0.2188. In the agentic setting, the Explorer Subagent plus File Budget configuration resolves 282/500 GPT-5.2 tasks (56.4%), 319/500 Grok 4.3 tasks (63.8%), and 148/500 Grok 4.1 Fast tasks (29.6%). Its mean cloud-facing file cost is 2.46, 0.56, and 0.99, respectively. Collateral leakage exceeds the price-matched random control by 2.20–2.66 percentage points for the per-file estimator and 5.37–5.93 percentage points for the dictionary estimator. Conclusion: Local intent inference and dynamic access control reduce cloudfacing disclosure while retaining substantial repair utility. The one-shot and agentic results support complementary privacy mechanisms, but their different datasets, models, and utility measures make cross-workflow comparisons descriptive rather than causal.
DATE: September 8, Tuesday @ 14:30
Place: EA 516