MammoClaw: Towards Skill-Evolving Agent Harness for Breast Cancer Mammography Analysis

Deep Breast Workshop on AI and Imaging for
Diagnostic and Treatment Challenges in Breast Care (Deep Breath)

[MICCAI 2026]

Munich, Bavaria, Germany
TL;DR MammoClaw turns a frozen, general-purpose MLLM into an auditable mammography agent. It gathers evidence by retrieving task-specific knowledge, inspecting ROIs, comparing paired CC/MLO views and checking the contralateral breast, and it evolves by distilling its own failed trajectories into reusable skills, with no weight updates. Every tool call and piece of evidence behind a decision is recorded and can be audited, a step towards trustworthy breast imaging agents.
MammoClaw Framework

MammoClaw Framework. The top panel illustrates the agent orchestration: a frozen MLLM alternates between reasoning, calls to deterministic mammography tools, and observations before producing a final prediction. The bottom panel shows the offline skill-evolution workflow: failed trajectories on a labeled reference set are analyzed by a teacher LLM and distilled into reusable textual skills, which are retrieved and injected into the agent context in later runs, without updating any model weights.

Abstract

In this work, we explore MammoClaw, a training-free agent framework that leverages frozen MLLMs for mammography analysis. To support agentic investigation, we equip the agent with lightweight mammography-specific tools for targeted image analysis, including ROI, paired-view, and contralateral-breast examination. MammoClaw iteratively gathers evidence through these tools, while skill evolution enables non-parametric adaptation by transforming failed trajectories into reusable guidance for later runs. We evaluate the framework on BI-RADS assessment and breast density estimation tasks. In our experiments, we find that tools alone do not reliably improve performance, whereas evolved skills can improve tool-use behavior and performance in some settings. Beyond these results, MammoClaw enables transparent inspection of evidence acquisition, tool interactions, and failure modes, facilitating the analysis and auditing of agent behavior. We view this work as an exploratory study of training-free, self-evolving agentic approaches for mammography and hope it provides a concrete starting point for future work on mammography-specific tools and self-evolution mechanisms.

🔥 Highlights

  1. Framework. We develop MammoClaw, a training-free agent harness for exploring agentic mammography analysis with frozen MLLMs, deterministic mammography tools, and an automated offline skill-evolution workflow.
  2. Tool Suite. We develop an initial suite of lightweight, deterministic, and model-free mammography tools for retrieving task knowledge, inspecting regions of interest, comparing paired and contralateral views, and gathering structured evidence, without relying on auxiliary segmentation or detection networks.
  3. Skill Evolution. We adapt an offline skill-evolution workflow that distills reusable reasoning guidance from failed trajectories and retrieves these skills in subsequent cases, allowing us to study experience-driven adaptation without updating the backbone model weights.
  4. Evaluation. We empirically investigate MammoClaw on BI-RADS assessment and breast density estimation. Tool augmentation alone does not necessarily improve performance (BI-RADS macro-F1 0.108 → 0.106), while skill evolution improves both performance (→ 0.148) and tool-use behavior.

🦞 MammoClaw Method Overview

Step 1: Agent Orchestration

Following the ReAct framework, the orchestration layer alternates between reasoning, tool invocation, and observation before producing a final prediction. A mammography case is represented as x = (I, q, 𝒞), where I is the input mammogram, q the task-specific question, and 𝒞 the candidate answer set. Let fθ be the frozen MLLM, 𝒜 the mammography tool library, and 𝒮 the skill bank. Before the first reasoning step, a lightweight retriever selects relevant skills, which are prepended to the initial prompt context ℋ0:

Sx = R(q, 𝒮)

At step t, the MLLM either produces a final prediction ŷ ∈ 𝒞 or selects a tool call ut = (At, pt), where pt holds tool-specific inputs such as ROI coordinates. The tool returns an observation that is appended to the context:

ot = At(x, pt),     ℋt = ℋt−1 ⊕ (ut, ot)

The resulting trajectory τ = {(ℋt−1, ut, ot)}t=1..k is an explicit record of the agent's reasoning and tool interactions, and is later used for skill evolution.

Step 2: Mammography Tools

The agent has access to a lightweight suite of deterministic, model-free tools: task-context tools, single-image ROI inspection, cross-view and contralateral comparison, and simple image-processing utilities (see Tool Suite). They return either text (BI-RADS definitions, domain knowledge, metadata) or images (ROI crops, paired-view composites, contralateral comparisons). Access alone does not guarantee effective use: the agent must still decide when to call a tool, which one, and how to interpret its output.

Step 3: Collecting Failed Trajectories

Skill evolution runs offline on a labeled reference set 𝒟ref = {(xj, yj)}j=1..N, held out from test evaluation. Starting from an empty skill bank 𝒮(0), at round r the agent is run on the reference set and failed trajectories are collected:

𝒯fail(r) = { (τj(r), xj, yj)  |  ŷj(r) ≠ yj }

Step 4: Distilling Skills

A teacher model G analyzes all failed trajectories together and proposes candidate skills 𝒮̃(r) = G(𝒯fail(r)). Skills capture reusable reasoning strategies, tool-use patterns, and corrections for recurring failure modes. They are not trained parameters: they are textual policies mined from prior failures and injected into the agent context at inference time.

Step 5: Updating the Skill Bank

Near-duplicate candidates are filtered using name-token Jaccard similarity and semantic cosine similarity over skill embeddings. The remaining novel skills Δ𝒮(r) are added to the bank:

𝒮(r+1) = 𝒮(r) ∪ Δ𝒮(r)

🛠️ Mammography Tool Suite

All tools are deterministic and model-free. Below is one representative call per tool, drawn directly from logged agent trajectories.

Task context
inspect_current_example

Returns the current example metadata.

Output: laterality=R, view=MLO, subject_age=61.0, has_paired_view=true, paired_view=CC

Task context
retrieve_knowledge

Retrieves task-specific domain knowledge, such as BI-RADS category definitions (0–6, including 4A/4B/4C) for BI-RADS assessment, or density-band definitions for density assessment.

Single-image inspection
inspect_mammogram_roi

Crops and enlarges a region using normalized 0–1000 image-grid coordinates, with optional contrast enhancement.

Input: bbox=[200, 200, 800, 800], enhance_contrast=true

ROI inspection example
Single-image measurement
measure_finding_size

Measures the width and height of a finding by drawing annotated measurement bars on a contextual crop.

Input: bbox=[0, 200, 250, 550], “spiculated mass at nipple/areolar region”

Finding size measurement example
Cross-view comparison
inspect_paired_mammogram_view

Returns the paired CC and MLO projections of the same breast side by side.

Paired view example
Cross-breast comparison
inspect_contralateral_breast

Compares the current breast with the contralateral breast from the same subject and view projection.

Contralateral breast example
Single-image density
estimate_breast_density

Estimates fibroglandular density via dual Otsu thresholding, e.g. density_percent=35.5, with a caveat not to assign A/B/C/D from the percentage alone.

Single-image quality
measure_image_sharpness

Scores technical sharpness via Laplacian variance and Tenengrad gradient energy, returning a sharpness rating.

Single-image texture
compute_tissue_statistics

Computes intensity mean, standard deviation, entropy, and skewness over the breast foreground (background excluded via Otsu).

🎯 Experiments

Datasets. We evaluate on Mammo-Bench for two tasks. Following its protocol, each source dataset is randomly split into 80% training and 20% evaluation. BI-RADS assessment uses KAU-BCMD (448 evaluation exams; BI-RADS 1/3/4/5 with 367/59/18/4 examples), which provides paired CC and MLO views and contralateral images. Breast density assessment uses DMID (108 evaluation exams; density A/B/C/D with 18/39/42/9 examples), which contains single-view mammograms only.

Agent and teacher. The agent backbone is Qwen3.5-35B-A3B. Skill evolution uses DeepSeek-V4-Flash as the teacher LLM, with one evolution iteration and at most 10 new skills per iteration. The reference set consists of 100 randomly sampled training examples with ground-truth labels. At inference, all task-relevant skills in the bank are injected into the agent context.

Metrics. We report macro-F1 with 95% bootstrap confidence intervals. Significance is assessed with paired bootstrap resampling (5,000 resamples) on macro-F1 and McNemar's test on accuracy, both Holm–Bonferroni corrected across the three pairwise comparisons per task.

Tools Alone vs. Evolved Skills

Adding the tool suite to the frozen agent leaves BI-RADS macro-F1 essentially unchanged (0.108 → 0.106, not significant). Adding evolved skills on top of tools raises macro-F1 to 0.148, a statistically significant gain over both the no-tool and tools-only settings (paired bootstrap, Holm-corrected p < 0.001). The benefit comes not from tool access alone, but from the guidance provided by the evolved skills.

Table: Effect of tools and evolved skills on BI-RADS assessment (KAU-BCMD).

MethodToolsSkillsMacro-F1 [95% CI]
Baseline✗✗0.108 [0.103, 0.113]
MammoClaw✓✗0.106 [0.093, 0.122]
MammoClaw✓✓0.148 [0.121, 0.180]

Statistical Significance Across Both Tasks

Both tasks show a similar trend: skill evolution improves macro-F1, whereas adding tools alone does not lead to a significant change. On breast density (n = 108), skills significantly improve over the tools-only setting (bootstrap p = 0.017), while the comparison with the no-tool baseline is not significant (p = 0.18).

Table: Macro-F1 differences between configurations, with Holm-corrected paired-bootstrap and McNemar p-values.

TaskComparisonΔMacro-F1Bootstrap pMcNemar p
BI-RADSNo tools → Tools−0.0020.700.0017
BI-RADSNo tools → Skills+0.040<0.001<0.001
BI-RADSTools → Skills+0.043<0.001<0.001
DensityNo tools → Tools+0.0190.700.062
DensityNo tools → Skills+0.0930.18<0.001
DensityTools → Skills+0.0740.0170.062

How Evolved Skills Change Tool Use

The evolved skills capture reusable reasoning strategies rather than explicit tool-selection rules: systematic inspection of suspicious regions, cross-view verification, sequential resolution of uncertainty, and grounding conclusions in tool observations. After skill evolution, the agent calls ROI inspection and contralateral comparison more often and the knowledge-retrieval tool less often, suggesting the skills supply some of the task guidance the agent would otherwise look up. Agents also make more tool calls per case overall. How tool usage relates to prediction quality remains an open question. See the full list of evolved skills below.

Tool calls with vs. without skills
Figure: Skill evolution increases evidence gathering per case. After skill evolution, the agent makes more tool calls on the BI-RADS assessment task, suggesting a more deliberate multi-step inspection process before final prediction.
Tool-call count distribution with vs. without skills
Figure: Distribution of tool calls per case before and after skill evolution on the BI-RADS assessment task.

Evolved Skills

These are the skills for the BI-RADS task, evolved automatically with DeepSeek-V4-Flash as the teacher model. They are distilled from failed reasoning trajectories on the reference set and stored in a reusable skill library. Rather than encoding explicit tool-selection policies, they capture reusable reasoning strategies: systematic evidence gathering, multi-view verification, uncertainty resolution, BI-RADS calibration, and grounding conclusions in tool observations. At inference, they are retrieved to guide the agent's reasoning without updating the underlying MLLM. Click a skill to expand it.

1 Verify Suspicious Findings with Multi-view Comparison
name: verify-suspicious-findings-with-multi-view-comparison  ·  category: general  ·  task: BI-RADS

Before concluding a suspicious finding like spiculation or architectural distortion, always compare with the paired view and contralateral breast to rule out artifacts or normal tissue overlap.

  1. Immediately invoke inspect_paired_mammogram_view to see if the finding is visible in both CC and MLO projections. If it is not clearly present in both, it may be a summation artifact.
  2. Then invoke inspect_contralateral_breast to check for bilateral symmetry. A finding that appears in the same location on the opposite breast is likely a normal anatomic variant.
  3. Only after confirming the finding is unilateral and present on two orthogonal views should you assign a BI-RADS category of 4 or 5.
  4. If the finding disappears or appears symmetric, reassess as benign or normal (BI-RADS 1 or 2).

Anti-pattern

  • Labeling a single-view finding as suspicious without multi-view confirmation.
  • Skipping contralateral comparison when suspecting pathology.
2 Systematic Quadrant Inspection for Subtle Findings
name: systematic-quadrant-inspection-for-subtle-findings  ·  category: general  ·  task: BI-RADS

When no obvious mass or calcifications are seen, systematically inspect all quadrants using ROI tools to avoid missing subtle asymmetries or microcalcifications.

  1. Divide the breast into quadrants mentally (upper outer, upper inner, lower outer, lower inner) or use the nipple as reference.
  2. For each quadrant, use inspect_mammogram_roi to zoom in, especially in areas where density appears slightly higher or where the tissue pattern changes.
  3. Look for clustered calcifications, subtle asymmetries, or architectural distortion that may be obscured by dense tissue.
  4. Compare the same quadrant in the contralateral breast to decide if a density is focal asymmetry or normal variant.
  5. Document any finding even if it seems probably benign; do not dismiss as normal without visual confirmation.

Anti-pattern

  • Stopping after a global impression without quadrant-level inspection.
  • Assuming dense tissue is normal without systematic checking; dense breasts can hide lesions, and systematic inspection reduces undercalling.
3 Tool Execution Before Conclusion
name: tool-execution-before-conclusion  ·  category: failure_recovery  ·  task: BI-RADS

After forming a hypothesis about a finding, immediately call the appropriate tool instead of describing a plan. Never present a tool plan without executing it.

  1. Formulate the specific question (e.g., “Is this density visible on the other view?” or “What are the margins of this mass?”).
  2. Immediately call the relevant tool—inspect_mammogram_roi, inspect_paired_mammogram_view, or inspect_contralateral_breast—without describing the plan in natural language.
  3. Wait for the tool output before refining your hypothesis. Do not characterize the finding based on the raw image alone.
  4. After the tool returns, reason from its output. Repeat if necessary but avoid redundant calls.

Anti-pattern

  • Writing sentences such as “I will now call inspect_mammogram_roi” without actually invoking the tool. The model must call the tool, not plan to call it.
4 Calibrate BI-RADS Using Feature Checklist
name: calibrate-birads-using-feature-checklist  ·  category: general  ·  task: BI-RADS

Use a systematic feature checklist to avoid extreme BI-RADS assignments (1 or 5) when intermediate categories (2, 3, 4) are more appropriate.

  1. BI-RADS 1 (Negative): No findings on any view. All tool outputs confirm absence of masses, calcifications, asymmetries, or distortion. Both contralateral and paired views show normal symmetric tissue.
  2. BI-RADS 2 (Benign): Clearly benign findings: popcorn calcifications, vascular calcifications, skin calcifications, well-circumscribed round masses with fat density (e.g., oil cysts, hamartomas).
  3. BI-RADS 3 (Probably Benign): New, solitary, well-circumscribed solid mass; focal asymmetry that is not changing; grouped punctate calcifications; mild architectural distortion not meeting spiculation criteria. Short-interval follow-up recommended.
  4. BI-RADS 4 (Suspicious): Indeterminate findings that do not have classic benign features: irregular margins, suspicious calcifications (pleomorphic, linear), new or evolving asymmetry with borderline features.
  5. BI-RADS 5 (Highly Suggestive of Malignancy): Classic malignant features confirmed on multiple views: spiculated mass, coarse heterogeneous calcifications with linear distribution, architectural distortion with retraction.

Anti-pattern

  • Jumping from BI-RADS 1 to 5 or 5 to 1 without intermediate evidence.
  • Failing to use the checklist to justify the exact category assigned.
5 Rule Out Artifact with Contralateral Comparison
name: rule-out-artifact-with-contralateral-comparison  ·  category: general  ·  task: BI-RADS

Before labeling a finding as suspicious, compare with the contralateral breast to identify normal variants that mimic pathology.

  1. Invoke inspect_contralateral_breast to view the mirror-image region of the opposite breast.
  2. If a similar density, pattern, or architectural appearance is present in the same location on the opposite side, it is highly likely a normal variant (e.g., asymmetric fibroglandular tissue, inframammary fold).
  3. If the finding is absent on the contralateral side, then proceed with further characterization.
  4. Document the comparison result in your reasoning. Bilateral symmetry strongly supports a benign or normal classification (BI-RADS 1 or 2).

Anti-pattern

  • Concluding a finding is suspicious without checking the opposite breast. Many normal variants are bilateral and symmetric.
6 Distinguish Benign from Probably Benign Calcifications
name: distinguish-benign-from-probably-benign-calcifications  ·  category: general  ·  task: BI-RADS

When evaluating calcifications, differentiate clearly benign patterns (BI-RADS 2) from probably benign patterns (BI-RADS 3) using morphology and distribution.

  1. Use inspect_mammogram_roi to zoom in on the calcifications and assess morphology at high resolution.
  2. Benign (BI-RADS 2): Popcorn (fibroadenoma), coarse (vascular), round/punctate scattered, skin calcifications, milk of calcium, dystrophic, suture. Typically large, well-defined, and not clustered.
  3. Probably Benign (BI-RADS 3): Grouped fine punctate (5+ in cluster), clustered but monomorphic, small round/oval in a cluster, no pleomorphism or linear shapes. Short-interval follow-up typical.
  4. Suspicious (BI-RADS 4/5): Pleomorphic, amorphous, fine linear/branching, coarse heterogeneous, with ductal distribution or segmental.
  5. Always confirm distribution on both views using inspect_paired_mammogram_view.

Anti-pattern

  • Assigning BI-RADS 2 to a cluster of fine punctate calcifications unless they are clearly scattered and not grouped.
  • Jumping to BI-RADS 4 for grouped punctate calcifications without considering BI-RADS 3.
7 Avoid Characterizing Findings Without Tool Confirmation
name: avoid-characterizing-findings-without-tool-confirmation  ·  category: failure_recovery  ·  task: BI-RADS

Never describe a finding as ‘spiculated’, ‘mass’, or ‘architectural distortion’ before using a tool to confirm its appearance.

  1. Formulate only as a hypothesis (e.g., “I see a region of increased density that might be a mass”).
  2. Immediately call the appropriate tool (inspect_mammogram_roi, inspect_paired_mammogram_view, or inspect_contralateral_breast) to gather evidence.
  3. After receiving tool output, use the observed features to characterize the finding. Describe what the tool shows, not what you think you see in the raw image.
  4. If the tool output does not clearly show the feature, do not assert it in reasoning.

Anti-pattern

  • Stating “The mammogram reveals a spiculated mass” before any tool call. Such statements are hallucinations and lead to misclassification.
8 Structured Reasoning Grounded in Tool Outputs
name: structured-reasoning-grounded-in-tool-outputs  ·  category: failure_recovery  ·  task: BI-RADS

Ensure every statement in the reasoning summary is directly supported by a tool finding listed in tool_findings, with explicit citation.

  1. List every tool call in tool_findings in the exact order they were made, with one bullet per tool.
  2. In reasoning_summary, for each claim about the image, explicitly reference which tool provided the evidence (e.g., “inspect_paired_mammogram_view confirmed the finding is present on both views”).
  3. Do not include any visual observations that were not obtained from a tool output. If an observation was made from the raw image and then confirmed with a tool, only the tool output counts as evidence.
  4. If multiple tools were used, explain how each tool contributed to the final decision.

Anti-pattern

  • Writing a reasoning summary that contains descriptions of findings without attributing them to specific tool calls. Every piece of evidence must be traceable to a tool output.
9 Use Uncertainty Resolution Sequence
name: use-uncertainty-resolution-sequence  ·  category: general  ·  task: BI-RADS

When uncertain about a finding, follow a structured sequence: inspect ROI, check paired view, compare contralateral, and optionally retrieve knowledge before finalizing.

  1. Step 1: Use inspect_mammogram_roi (up to 2 times) to zoom in on the area and characterize margins, shape, density, and internal features.
  2. Step 2: Use inspect_paired_mammogram_view to confirm the finding appears in the second projection. If it does not, it is likely superimposition.
  3. Step 3: Use inspect_contralateral_breast to check if the finding is bilateral. If symmetric, it is a normal variant.
  4. Step 4 (if needed): Use retrieve_knowledge to resolve definitional uncertainty about categories (e.g., what exactly constitutes a probably benign calcification).
  5. Only after completing this sequence should you assign a BI-RADS category.

Anti-pattern

  • Making a final decision while still uncertain. The sequence is designed to reduce uncertainty incrementally.
  • Skipping steps, which leads to over- or under-calling.

Agent Trajectory: A Suspicious BI-RADS 4 Case

Browse all trajectories

We provide full agent trajectories for 100 examples, with every reasoning step, tool call and tool output (large page, ~60 MB)

A representative trajectory in which the tool-augmented agent with evolved skills reaches a correct suspicious assessment. An ROI inspection localizes architectural distortion, paired-view comparison verifies that the finding persists across projections, and contralateral comparison rules out a symmetric normal variant. Every tool interaction and piece of intermediate evidence is exposed, which makes agent behavior auditable.

Retrieved skills: verify-suspicious-findings-with-multi-view-comparison rule-out-artifact-with-contralateral-comparison structured-reasoning-grounded-in-tool-outputs calibrate-birads-using-feature-checklist

User query 💡 LLM reasoning Tool call Final decision

User Query
“What is the correct BI-RADS value for this case? Choose exactly one: 0.0–6.0.”
Input mammogram
💡 Step 1 · Initial Reasoning
“The user wants me to determine the BI-RADS category for this mammogram. Let me start by examining the image and then use the appropriate tools to make a thorough assessment. I should first inspect the current example to understand the metadata, then systematically evaluate the breast.”
Step 1 · Tool Call: inspect_current_example()
Output: source_dataset=kau-bcmd, laterality=R, view=MLO, subject_age=45.0, has_paired_view=true, paired_view=CC
💡 Step 2 · Reasoning (before ROI call)
“I see an MLO view of the right breast. There appears to be an area of architectural distortion or asymmetry in the lower portion of the breast. Let me first inspect this area more closely with an ROI, then check the paired view and contralateral breast.”
Step 2 · Tool Call: inspect_mammogram_roi(bbox=[200,650,450,950])
Input: description=“area of architectural distortion/asymmetry in lower breast”
The contrast-enhanced crop shows radiating lines converging toward a central point.
ROI crop
💡 Step 3 · Reasoning (after ROI, before paired-view call)
“The ROI shows an area of architectural distortion with spiculations radiating from a central point near the skin/chest wall area. This looks suspicious. I need to check the paired view to confirm if this is a real finding or an artifact.”
Step 3 · Tool Call: inspect_paired_mammogram_view()
Side-by-side panel, current R MLO (left) vs. paired R CC (right).
The finding is present on both R MLO and R CC, supporting a real 3D finding rather than a summation artifact.
Paired view
💡 Step 4 · Reasoning (after paired view, before contralateral call)
“The paired view shows the architectural distortion/linear density in the lower posterior aspect of the breast on both the MLO and CC views. This confirms it is a real finding. Now I need to check the contralateral breast to see whether it is a normal bilateral variant.”
Step 4 · Tool Call: inspect_contralateral_breast()
Side-by-side panel, current R MLO (left) vs. contralateral L MLO (right).
The left breast lacks a corresponding focal distortion, supporting a unilateral suspicious finding.
Contralateral comparison
Step 5 · Final Decision ✓ (matches ground truth BI-RADS 4)
“The mammogram reveals a unilateral area of architectural distortion visible on both MLO and CC views, confirmed by multi-view comparison. The contralateral breast lacks a corresponding finding, indicating it is not a normal variant. Architectural distortion is a suspicious finding that typically warrants a BI-RADS 4 assessment.” ⇒ BI-RADS 4

❗ Limitations

  • Outcome-based, not process-based, evaluation. Our metrics (macro-F1, significance tests) only score the final prediction. They do not assess whether the agent's reasoning, tool use, and gathered evidence were clinically sound, so a correct answer may rest on flawed reasoning and vice versa. Step-level and trajectory-level metrics, ideally with expert review, are needed to evaluate the process itself.
  • Evaluation covers two tasks, a single frozen backbone (Qwen3.5-35B-A3B, which is modest in the no-tools setting), and one dataset per task.
  • Evolved skills come from a labeled reference set and have not been validated by clinical experts. Their clinical relevance, robustness, and transfer under dataset shift remain open.
  • The effects of the system prompt, tool descriptions, teacher LLM, skill-generation prompt, and retrieval mechanism have not been ablated.
  • Skill evolution uses a single iteration and a 100-example reference set; multi-round evolution, skill pruning, and cross-task transfer are left to future work.

🤝 Open to Collaboration

LOOKING FOR COLLABORATORS Because of limited compute and resources, the skill evolution reported in this paper uses only a single evolution round on a 100-example reference set. The natural next step, multi-round skill evolution, is left unexplored, and it is the direction I am most interested in pursuing with collaborators.
  • Multi-round skill evolution. In the current setup, failed trajectories are collected once, a teacher model distills them into skills, and those skills are used for all later runs. How this behaves when the loop is repeated many times, whether performance improves steadily or saturates, and how to validate and prune a growing skill bank, is what I would most like to work on.
  • Feedback on the evolved skills. If you are a radiologist or clinical researcher, I would value your view on whether the evolved skills encode clinically meaningful reasoning rather than dataset-specific shortcuts, and where the agent's reasoning in the logged trajectories is unsound.

Feedback and collaboration opportunities are welcome by email.

Citation


  @InProceedings{Nakka_2026_MICCAI,
    author    = {Nakka, Krishna Kanth},
    title     = {MammoClaw: Towards Skill-Evolving Agent Harness for Breast Cancer Mammography Analysis},
    booktitle = {Proceedings of the Deep Breast Workshop on AI and Imaging for Diagnostic and Treatment Challenges in Breast Care, MICCAI 2026},
    year      = {2026},
}
  

Visitors

Acknowledgement

This website is adapted from Nerfies, licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.