25 points nateb2022 22 hours ago 7 comments

iAMkenough 1 hour ago | parent

So that’s why big ballz (co author of the OP) put our PII in an insecure AWS instance via DOGE’s starlink terminal

https://www.csoonline.com/article/4046997/whistleblower-doge...

sgnelson 57 minutes ago | parent

The skeptic in me really can't trust our current government to protect my information.

BowBun 53 minutes ago | parent

Good thing you don't need much trust in this case. Source available here - https://github.com/nationaldesignstudio/rampart

I suppose the model could be doing stuff, but you can also switch that out for your own with the source.

goodmythical 42 minutes ago | parent

I mean, it can be run locally, so you don't necessarily have to trust it, unless there's been any model-as-a-vector CVE.

That hasn't happened yet, has it? Where running a model from HF directly compromises the machine as opposed to some breakout or exfil done by the model after the fact?

That said, if you're filing your taxes or have a 'real' ID, the government already has all pertinent information required to fuck you over either deliberately or via a leak, so...

dwa3592 41 minutes ago | parent

I have worked in this field and I am the author of this package - https://github.com/deepanwadhwa/zink

A few things jump out since this is done by the government:

- the lowest hanging fruit for this problem is to clearly tell people (citizens) not to share any personal info with chatbots which can cause financial harm or identity theft. the example on the page shows a person sharing their SNN with a chatbot to help them find an apartment - "My name is Maria Garcia, my Social Security number is 123-45-6789, and I make $1,950 a month. Can you help me find affordable housing?" - why?? this is the opposite of what i would expect a government to advise their citizens.

- it's never too late for a good policy; the government should have extended HIPPA and other data privacy laws to AI companies - the AI company must not store anyone's SSN, no matter how stupid the user is. It should be on the AI company to not store it; so this type of layer should be on the AI company's side.

- technical; there are quasi identifiers of privacy (that's what my package targets) that are asymptotically hard to to deal with - meaning - if you remove everything that can leak your privacy the text would become meaningless. i don't think rampart can solve for that either and it should be clearly said on the website.

bob1029 40 minutes ago | parent

I have presented approaches like this to banking clients and they are still not very interested. The only thing that makes these people happy is zero data retention and deterministic redaction at the source. Regex over arbitrary string literals does not represent determinism in this context.

If your product is handling natural language conversations from end customers, there is not much you can do to prevent the occasional PII leak without ruining the rest of the pie. ZDR is your best mitigation if you actually want the magical AI experience to work the way the investors hope it can.

PII can often become disclosed by way of many correlated factors that are not considered PII on their own. Even a perfect AI system cannot capture all of these relationships. You could probably locate where I live within a 20 mile radius if you spent enough time analyzing my HN comments over the years. Not one of these comments on their own would trigger a PII filter.

nhinck2 38 minutes ago | parent

98.4% is nowhere near good enough to call it PII redaction.