FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios

arXiv:2605.00706v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied in financial scenarios. However, they may produce harmful outputs, including facilitating illegal activities or unethical behavior, posing serious compliance risks. To systematically evaluate LLM safety in finance, we propose FinSafetyBench, a bilingual (English-Chinese) red-teaming benchmark designed to test an LLM's refusal of requests that violate financial compliance. Grounded in real-world…

cs.CL updates on arXiv.org · May 4 · 1 min read

From the source

Originally published at cs.CL updates on arXiv.orgRead at source →