SmarTag
News labels, relevance, sentiment, summaries (9 card tables).
Buy by the field
23 priced fields
Pick the fields you need — priced individually below.
20% off buying every field
Pay annually · save 20% — $21,792/yr
ⓘ Take the whole set from the field list below — select everything, add it to your basket, and check out there.
SmarTag
smartagSource DataNews labels, relevance, sentiment, summaries (9 card tables).
Daily source stitch · data frontier 2026-08-03
9. Field reference — every column
The full field dictionary — every table, every column, with its type and meaning. One real sample row is shown per table.
news_info
Table · pick one of 9
news_info
14 columns · grain (primary key): newsid
| Column | Name | Type | Null | Description |
|---|---|---|---|---|
operation | Change-feed flag (A/U/D) | string | yes | Change-feed flag marking whether the row adds (A), updates (U), or deletes (D) an article record, so subscribers can apply the data as an incremental change stream. SmarTag news is append-only, so every row in the dataset is an Add (A). |
newsid ·PK | News article ID | string | no | The unique identifier of the article (a numeric article id) and the primary key of this table — one row per news article. Every tag table references it, so it is the join key from an article to its companies, people, products, regions, concepts, events, industries and sentiment. |
newstitle | Article headline | string | yes | The article's headline as captured. Populated on essentially every article. |
newstitle_cn | Article headline (Chinese) | string | yes | The article headline in Chinese. For Chinese-language sources it matches the headline; it carries the Chinese rendering when the original differs. Populated on essentially every article. |
newsts | News timestamp (capture time) | timestamp | no | The time the article was captured and tagged — the point-in-time anchor for as-of and look-ahead-safe queries. Values span 2017-01-01 through the current day; date each tag by this field to reconstruct exactly what was knowable on any past date. |
newsoriginalts | Original publication time | timestamp | yes | The article's original publication time as reported by the source — as opposed to the capture time (News timestamp). Spans 2017-01-01 onward and is populated on essentially every article; use it when you need the true publication moment rather than when SmarTag ingested the piece. |
newsurl | Article URL | string | yes | The web address the article was published at. Populated on essentially every article. |
newssource | News source / publisher | string | yes | The publisher or portal the article came from — e.g. 新浪网 (Sina), 东方财富网 (Eastmoney), 同花顺财经 (Flush), 证券之星, 金融界. Thousands of distinct sources appear; about 5.8% of rows are blank. |
newssummary | Article summary | string | yes | A short summary or lead paragraph of the article. Present on essentially every article (under 0.001% blank). This is a summary/abstract — the full article body is a separate product, not part of SmarTag. |
whitelistflag | Whitelisted flag | integer | yes | A flag marking whether the article is on SmarTag's curated whitelist of relevant financial-news items: 1 = whitelisted (~90% of articles), 0 = not (~10%). Filter to 1 for the cleaned financial-news stream. |
newsextid | External article ID | string | yes | An external source identifier (a hash) for the article, used to de-duplicate against the source feed; about 5.4% of rows are blank. |
emotionindicator | Sentiment label (0/1/2) | integer | yes | The article-level sentiment: 0 = neutral, 1 = positive, 2 = NEGATIVE. Note the mapping — 2, not 0, is negative. Across the full history about 52% of articles are positive, 24% negative and 24% neutral. |
emotionweight | Sentiment confidence weight | number | yes | The confidence weight of the sentiment label, ranging 0.33–1.0 (median ~0.88) — higher means the model is more certain of the call. Pair it with the sentiment label to drop low-confidence classifications. |
emotiondetail | Sentiment class weights | string | yes | The full sentiment breakdown — the model's weight on each of the three classes, formatted {0=neutral, 1=positive, 2=negative}, e.g. {0=0.0088, 1=0.9892, 2=0.002}. The Sentiment label is the arg-max of these three weights; use this field when you want the soft scores rather than the hard label. |
Sample row:
operation: A
newsid: 136341720
newstitle: 中国学者首次获得门捷列夫国际基础科学奖
newstitle_cn: 中国学者首次获得门捷列夫国际基础科学奖
newsts: 2026-07-13 00:00:06
newsoriginalts: 2026-07-12 23:53:00
newsurl: http://www.chinanews.com.cn/sh/2026/07-12/10658059.shtml
newssource: 中国新闻网
newssummary: 中新社合肥7月12日电 (记者 吴兰)记者12日从中国科学技术大学获悉,联合国教科文组织近日正式宣布,中国科学院院士、中国科学技术大学教授潘建伟荣获第三届门捷列夫国际基础科学奖…
whitelistflag: 1
newsextid: b300f89643ec13e336173e8a4ad9cfeb
emotionindicator: 1
emotionweight: 0.99
emotiondetail: {0=0.0088, 1=0.9892, 2=0.002}
Sourced as-delivered from ChinaScope, with resolved codes and structural links computed by Numinor. Licensing passes through to you; fields are served exactly as the vendor issues them.
Access & FAQ
- How do I buy just a few fields?
- Check the fields you want and you'll see your running subtotal. Select every field and the whole-set price applies automatically. Add them to your basket and check out — this dataset, or fields combined across several.
- Can I evaluate these fields in the Matrix before I buy?
- Yes — load the actual source tables in the Matrix and analyze them with an AI agent (point-in-time lagged) to check the fields fit your use case before you buy. Once your selection is activated, export the fields you choose via the API.
- How is it delivered?
- Apache Parquet over a signed-URL REST API, the same pipe as Construct Data. The free Matrix sandbox serves a time-delayed view for evaluation.
- When is each update available, and how reliable is it?
- Source data refreshes on the ChinaScope cadence and is delivered through the same signed-URL pipeline as Construct Data, with the same manifest/status behaviour. Evaluate freshness yourself in the Matrix before you commit.
- Why are some fields free and others cost thousands?
- Price tracks non-replicability. Resolved codes and structural links are the moat; machine scores and raw text are cheap; identifier keys are free. Figures are priced per table; codes are charged once per dataset.
Still have questions? Contact Sales →