Beyond Citation Blog

Emily Huang is a research scientist at Google

One problem keeps showing up in database work: a name can point to more than one person. Emily Huang is a good example. Public records and profile pages surface several Emilys with research roles, but the Google-linked entry is thin, and that thinness matters because it limits what a searcher can safely infer.

I spend a lot of time on that kind of problem. A database often gives a name first and a story later. If the story is incomplete, I have to slow down and separate the confirmed parts from the guessed parts. That is the work. It is not flashy, but it keeps a catalog record, a profile page, or a scholar listing from turning into a trap.

The phrase “research scientist at Google” sounds precise. It is precise only if the record behind it is tied to the right Emily Huang and if the date is clear. In practice, a name search can surface multiple people with the same name, plus pages that mix current and past roles. That is why I treat the job title as a claim to verify, not a fact to repeat carelessly.

What makes this harder is that academic identity now lives across several systems. A person may appear in a scholarly index, a university profile, a company page, or a social network. Each place may show a different slice of the same life. Some pages are curated. Some are stale. Some are only there because a publication index matched a name string.

The beginner mistake is to think that a name search is a person search. It is not. It is a string match. A string match can be useful, but it has weak grip. It needs context. Affiliation, coauthors, subject area, and publication dates all help tell one Emily Huang from another.

I teach this as a simple habit: do not stop at the first hit. Read the surrounding details. Look for institutional ties. Check whether the subject area fits the person you mean. Then ask a sharper question. Is this a confirmed identity, or only a plausible one?

Here is a small example. Suppose a search returns “Emily Huang” and one result says Google. Another result says a university post. A third result points to a medical author list. Those three records may describe three different people. They may also reflect one person whose role changed over time. The searcher cannot know that from the name alone. The useful move is to compare the field, the coauthors, and the dates before treating any result as the same person.

That example sounds modest because it is. But it is the sort of modesty databases require. A system that indexes names well still fails if the user reads too much into a bare match. This is one reason I keep insisting that discovery tools are only as honest as the fields behind them. When metadata is sparse, search looks confident while meaning stays fuzzy.

For a project like Beyond Citation, that matters in a very practical way. We are not trying to flatter databases. We are trying to describe how they behave. A name-based result may be enough to begin, but it is rarely enough to finish. The harder work is telling readers where a database helps and where it leaves room for error.

Emily Huang’s case also shows why authority control matters. That phrase sounds technical, but the idea is plain. Authority control means a system uses a stable form of a name, so one person is not split into many false identities. When that control is weak, the same scholar can appear in fragments. When it is stronger, the searcher gets a better chance of joining those fragments into one person.

Still, no system fixes everything. Even a well-kept profile can lag behind reality. A person may have changed jobs. A publication list may omit context. A company page may only show one role and hide the research history behind it. I have to say that plainly because the record does not become true just because a database displays it.

What I can do, and what this lesson is built to do, is show how to read the record with care. Start with the exact wording. Ask what kind of source it is. Check whether the detail comes from a profile, an index, or a scraped page. Then decide how much weight that detail deserves. That sequence is slow, but it is clean.

The practical payoff is simple. After this kind of reading, a user can tell the difference between a likely match and a secure identification. That sounds small. It is not. It keeps a citation trail, a faculty search, or a background check on the right track before confusion hardens into fact.

That is the sort of plain work The Source List is meant to support: one digital source worth knowing, one search tip, and one honest limitation.