feat(read_file): let capable models view images, video and PDFs
Nightly Build / build (push) Successful in 6m38s
Nightly Build / build (push) Successful in 6m38s
When the resolved model declares an input modality (vision → images,
video, document → PDFs), read_file now hands a binary media file back to
the model as native input instead of failing on non-UTF-8 bytes.
- ToolResult gains a Media { text, media } variant carrying MediaRef
{ host_path, mime }; the tool message keeps only the text note, the
bytes travel out of band in a new additive chat_llm_tools.media column
(mirrors preview_old/new).
- read_file sniffs the resolved host file; a recognized medium becomes a
Media result (with a neutral note), everything else keeps the textual
path. Capability gating lives in the message builder, so read_file
never needs the model caps and degrades cleanly on a text-only model.
- MessageBuilder inlines current-turn tool media as a synthetic user
message right after the tool-result group (media_turn_start boundary,
so older turns are never re-billed), reusing media.rs primitives via a
new inline_paths helper that contains against the caller's workspace
roots. OpenAI forwards the parts verbatim; the Anthropic client now
also translates the PDF `file` part into a native `document` block
(image_url → image was already handled).
- read_file's description is annotated per serving model in
call_llm_round, listing the formats it can open, so the model knows
reading one shows it the content.
Tests: media sniff (PDF), PDF file-part build, inline_paths containment
+ capability gating, capability hint, read_file media-vs-text, Anthropic
file→document, and the owner-schema-stands-alone check with the new
column.
This commit is contained in:
@@ -181,6 +181,11 @@ fn convert_user_content(content: &Value) -> Value {
|
||||
blocks.push(block);
|
||||
}
|
||||
}
|
||||
"file" => {
|
||||
if let Some(block) = parse_data_document(&p["file"]) {
|
||||
blocks.push(block);
|
||||
}
|
||||
}
|
||||
other => tracing::warn!(part_type = other, "dropping content part unsupported by Anthropic"),
|
||||
}
|
||||
}
|
||||
@@ -198,6 +203,18 @@ fn parse_data_image(image_url: &Value) -> Option<Value> {
|
||||
}))
|
||||
}
|
||||
|
||||
/// `{"file_data": "data:application/pdf;base64,<data>"}` → an Anthropic base64
|
||||
/// `document` block (the native PDF input). Only base64 data URLs are supported;
|
||||
/// the OpenAI `file` part is what the media pipeline emits for a PDF.
|
||||
fn parse_data_document(file: &Value) -> Option<Value> {
|
||||
let url = file["file_data"].as_str()?;
|
||||
let (mime, data) = url.strip_prefix("data:")?.split_once(";base64,")?;
|
||||
Some(json!({
|
||||
"type": "document",
|
||||
"source": { "type": "base64", "media_type": mime, "data": data },
|
||||
}))
|
||||
}
|
||||
|
||||
#[async_trait]
|
||||
impl ChatbotClient for AnthropicClient {
|
||||
async fn chat(
|
||||
@@ -435,4 +452,24 @@ mod tests {
|
||||
]));
|
||||
assert_eq!(v, json!([{ "type": "text", "text": "t" }]));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn user_content_file_part_becomes_document_block() {
|
||||
// The OpenAI `file` part (emitted by the media pipeline for a PDF) becomes
|
||||
// an Anthropic native `document` block.
|
||||
let v = convert_user_content(&json!([
|
||||
{ "type": "text", "text": "read this" },
|
||||
{ "type": "file", "file": { "filename": "a.pdf", "file_data": "data:application/pdf;base64,QUJD" } },
|
||||
]));
|
||||
assert_eq!(v, json!([
|
||||
{ "type": "text", "text": "read this" },
|
||||
{ "type": "document", "source": { "type": "base64", "media_type": "application/pdf", "data": "QUJD" } },
|
||||
]));
|
||||
|
||||
// A non-data file_data (or missing) is dropped, not forwarded.
|
||||
let v = convert_user_content(&json!([
|
||||
{ "type": "file", "file": { "filename": "a.pdf", "file_data": "https://example.com/a.pdf" } },
|
||||
]));
|
||||
assert_eq!(v, json!([]));
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user