Giskard
company
Verified
AI & ML interests
LLM, Agents, Evaluation, Evals, AI Quality, AI Security, Red-teaming
Recent Activity
View all activity
kevin-giskardย
updated a
dataset about 1 month ago
kevin-giskardย
published a
dataset about 1 month ago
kevin-giskardย
updated a
dataset about 1 month ago
kevin-giskardย
published 2
datasets about 2 months ago
pierljย
updated a
dataset 3 months ago
weixuan-giskardย
updated a
dataset 3 months ago
pierljย
published a
dataset 4 months ago
julien-cย
submitted a
paper to Daily Papers 6 months ago
davidberenstein1957ย
posted an update 8 months ago
Post
2872
๐จ Phare LLM benchmark V2: Reasoning models don't guarantee better security
Read the full blog here: https://huggingface.co/blog/davidberenstein1957/phare-llm-benchmark-v2
Read the full blog here: https://huggingface.co/blog/davidberenstein1957/phare-llm-benchmark-v2
pierljย
updated a
dataset 8 months ago
alexcombessieย
updated a
Space about 1 year ago
davidberenstein1957ย
posted an update about 1 year ago
Post
638
Announcing RealPerformance, a dataset of functional issues of language models that mirrors failure patterns identified through rigorous testing in real LLM agents
https://huggingface.co/blog/davidberenstein1957/realperformance-llm-business-compliance
https://huggingface.co/blog/davidberenstein1957/realperformance-llm-business-compliance
Update README.md
1
#3 opened about 1 year ago
by
davidberenstein1957
Update README.md
#2 opened about 1 year ago
by
davidberenstein1957
Update README.md
#2 opened about 1 year ago
by
davidberenstein1957
davidberenstein1957ย
posted an update about 1 year ago
Post
424
๐จ LLMs recognise bias but also reproduce harmful stereotypes: an analysis of bias in leading LLMs
I've written a new entry in our series on the Giskard, BPIFrance and Google Deepmind Phare benchmark(phare.giskard.ai).
This time it covers bias: https://huggingface.co/blog/davidberenstein1957/llms-recognise-bias-but-also-produce-stereotypes
Previous entry on hallucinations: https://huggingface.co/blog/davidberenstein1957/phare-analysis-of-hallucination-in-leading-llms
I've written a new entry in our series on the Giskard, BPIFrance and Google Deepmind Phare benchmark(phare.giskard.ai).
This time it covers bias: https://huggingface.co/blog/davidberenstein1957/llms-recognise-bias-but-also-produce-stereotypes
Previous entry on hallucinations: https://huggingface.co/blog/davidberenstein1957/phare-analysis-of-hallucination-in-leading-llms
davidberenstein1957ย
updated a
dataset about 1 year ago