Skip to content

Milestones

List view

  • A method for blind detection of ROME edits

    Overdue by 1 month(s)
    Due by June 30, 2026
  • Conduct the security landscape mapping for the ROME method. Covert most of the SotA open-source models and evaluate the ROME applicability and robustness.

    Overdue by 6 month(s)
    Due by January 31, 2026
    1/1 issues closed
  • Break the current SotA detection presented by Youssef et. al. The authors discovered that GPT2 models have significant pair-wise cosine distance introducted into the edited layer weights by the ROME. Get rid of this cosine distance during the delta optim.

    No due date
    2/2 issues closed
  • An auto selection of the memory layer based on the results of causal tracing. This functionality is crucial for achieving seamless ROME edits.

    Overdue by 8 month(s)
    Due by December 1, 2025
    3/3 issues closed