Fou tee -yea -old
Zey ep
Demi bas
has
fou d
that
ChatGPT-4o
is
less
accu ate
tha
a
me tal
health-focused
AI
model
a d
a
simple
machi e-lea i g
system
at
detecti g
st ess
i
huma
w itte
text.Zey ep,
a
eighth-g ade
stude t
at
T a sit
Middle
School
i
East
Amhe st,
New
Yo k,
tested
fou
diffe e t
models
usi g
mo e
tha
3,500
Reddit
posts.
Acco di g
to
Society
fo
Scie ce,
the
posts
had
al eady
bee
labelled
by
huma s
as
showi g
st ess
o
o
st ess.He
p oject,
titled
“Evaluati g
the
eliability
of
La ge
La guage
Models
fo
st ess
detectio ”,
has
ea ed
he
a
place
amo g
the
fi alists
i
the
2025
The mo
Fishe
Scie tific
Ju io
I ovato s
Challe ge.
The
competitio
ecog ises
you g
stude ts
wo ki g
o
scie ce-based
p ojects.Zey ep
became
i te ested
i
the
p oject
afte
speaki g
with
a
family
f ie d
who
is
a
psychologist.
The
psychologist
told
he
that
some
health
i su a ce
compa ies
we e
explo i g
la ge
la guage
models
(LLMs)
as
cheape ,
24/7
alte atives
to
huma
the apists.
Zey ep
wo de ed
whethe
AI
systems
could
actually
be
t usted
to
ide tify
st ess.
Testi g
AI
models
To
test
the
models,
Zey ep
used
a
dataset
called
D eaddit.
It
co tai s
3,553
Reddit
posts
that
huma
ate s
had
labelled
acco di g
to
whethe
they
co tai ed
sig s
of
st ess.
She
gave
the
data
to
fou
diffe e t
models
i cludi g
Bidi ectio al
E code
Rep ese tatio s
f om
T a sfo me s
(BERT),
Me talBERT,
Ra dom
Fo est
a d
ChatGPT-4o.
Me talBERT
is
a
ve sio
of
BERT
desig ed
fo
me tal
health- elated
la guage,
while
Ra dom
Fo est
is
a
basic
machi e-lea i g
tech ique
that
uses
multiple
decisio
t ees
to
make
p edictio s.Zey ep
asked
each
model
to
ide tify
which
posts
showed
st ess.
She
the
used
a
measu e
called
a
F1-sco e
to
compa e
thei
pe fo ma ce.
The
sco e
co side s
both
how
accu ately
a
model
ide tifies
st ess
a d
how
ofte
it
misses
st ess
o
w o gly
labels
a
post
as
showi g
st ess.Me talBERT
pe fo med
the
best
i
he
testi g,
with
a
sco e
of
about
82
pe ce t.
BERT
followed
with
about
79
pe ce t.
ChatGPT-4o
sco ed
about
74
pe ce t.
It
also
pe fo med
wo se
tha
the
Ra dom
Fo est
model,
which
was
i cluded
as
a
simple
baseli e
fo
compa iso .The
esult
su p ised
Zey ep
because
Ra dom
Fo est
is
a
much
simple
machi e-lea i g
method
a d
does
ot
u de sta d
la guage
a d
co text
i
the
same
way
a
LLM
does.
“ChatGPT
pe fo mi g
badly
was
‘ eally
su p isi g,’”
Zey ep
said.She
fou d
it
pa ticula ly
i te esti g
that
a
simple
model
could
outpe fo m
a
LLM
with
millio s
of
pa amete s.
“Ra dom-fo est
is
‘supposed
to
be
a
ve y
simple
a d
old
tech ique.
So
I
just
put
it
i
as
a
baseli e,’”
Zey ep
said,
as
quoted
by
Scie ce
News
Explo es.
“That
was
ve y
i te esti g;
how
somethi g
so
small
a d
simple
was
able
to
beat
a
LLM
like
ChatGPT
that
used
millio s
of
pa amete s
a d
had
so
much
codi g
go
i to
it,”
she
added.
What
esults
mea
Zey ep’s
fi di gs
made
he
questio
whethe
ge e al-pu pose
LLMs
a e
eliable
e ough
to
be
used
fo
me tal
health
assessme t.
A
la ge
la guage
model
is
a
type
of
machi e-lea i g
system
t ai ed
o
ve y
la ge
amou ts
of
text.
It
lea s
patte s
i
la guage
a d
uses
them
to
p oduce
espo ses.He
esults
led
he
to
co clude
that
LLMs
should
ot
eplace
huma
the apists.“My
p oject
shows
that
LLMs
a e
cu e tly
u eliable
a d
u safe
to
deploy
as
diag ostic
tools,”
Zey ep
said.She
added
that
the
fi di gs
did
ot
mea
that
LLMs
a e
bad
o
ca ot
be
useful.
But,
they
show
that
ge e al-pu pose
AI
systems
may
ot
be
suitable
fo
a
task
as
se sitive
as
assessi g
me tal
health.“We
should
be
mi dful
with
AI,
because
it
does ’t
eally
have
a
acceptable
g ade
i
me tal
health,”
Zey ep
said.
“That
does ’t
mea
that
LLMs
a e
bad,
because
they’ e
fo
ge e al
use.
They’ e
ot
ecessa ily
mea t
fo
me tal
health,”
she
added.Zey ep
also
suggested
that
LLMs
could
pote tially
have
a
diffe e t
ole.
I stead
of
eplaci g
me tal
health
p ofessio als,
they
might
help
ide tify
people
who
a e
st uggli g
a d
efe
them
to
a
me tal
health
p ofessio al.
Zey ep
wa ts
to
study
AI
bias
The
p oject
also
made
Zey ep
i te ested
i
whethe
LLMs
might
show
diffe e t
esults
depe di g
o
a
pe so ‘s
ge de .
“O e
way
I
feel
I
could
expa d
it
is
seei g
whethe
LLMs
ca y
biases
towa d
diffe e t
ge de s,”
Zey ep
said.She
said
she
had
ead
about
cases
whe e
docto s
dismiss
symptoms
epo ted
by
female
patie ts
because
they
believe
wome
a e
exagge ati g.
She
sees
this
as
a
example
of
pe so al
bias
a d
wa ts
to
k ow
whethe
AI
systems
could
show
simila
patte s.Si ce
LLMs
a e
t ai ed
usi g
la ge
amou ts
of
text
c eated
by
people,
Zey ep
believes
they
ca
also
pick
up
huma
biases.Zey ep’s
wo k
ea ed
he
a
fi alist
place
i
the
2025
The mo
Fishe
Scie tific
Ju io
I ovato s
Challe ge.
She
hopes
to
become
a
compute
scie tist.She
said
she
e joys
p og ammi g
but
is
pa ticula ly
i te ested
i
how
compute
scie ce
ca
be
applied
to
eal-wo ld
p oblems.
